Eh, I had to get to the jobs I didn’t really enjoy eventually.
I wound up at Oracle because I badly wanted to leave Amazon. Never the best reason for taking a role, but we’ve all been there. Also, true story, Oracle rejected me back in 2002 because I didn’t have a college degree, so I felt a little bit like I was getting some of my own back. Not really rational but we’re all human.
As I mentioned in an earlier post, about a month after I got to Oracle my hiring manager’s boss mentioned to me that he didn’t really think my role made sense, which was very alarming. Let’s drill into that one.
I was hired as an SRE manager, with four or so SREs reporting into me. This was basically a lateral move from Amazon. In theory my SREs were going to embed into engineering teams and provide SRE skills where needed. In practice, Oracle Cloud was following the Amazon model where engineering teams were expected to own their code in production, meaning they’d best have those SRE skills already. And since most of the initial leadership team was from Amazon, all those core teams had been built with that model in mind.
I’d assumed that my first couple of months were going to be spent moving my team off the work they were busy with and then inserting them into teams who would be happy to have the help. This was not at all true; nobody had any pre-existing interest in my team’s services. Ooops.
So OK. This was clearly going to require some effort, and the work my team was doing at the time wasn’t going to wrap up trivially; I had some time to solve this.
I spent that time meeting people and working my connections. I did have one buddy at Oracle, Dalibor, who’d been a peer of mine at Amazon and was now a director running the object storage team. I talked him into taking on one of my people. Long story short, both because it’s a long one and because I don’t remember all the details, I managed to find reasonable homes for the rest of them. It was still a bad structure: organizationally, the teams making use of my services would have been better off hiring their own permanent SREs. It was better than nothing, though. I felt a bit at loose ends myself, but I figured I could make myself useful by looking for common problems and coming up with solutions.
Somewhere in there my hiring manager got moved into an IC role, and I wound up reporting into Karl, my hiring manager’s peer, who was sympathetic. I liked working for Karl – he had a good sense of the politics of the place and cared about my career. Through no fault of his own, he started me on the real roller coaster ride by handing me his old operational business intelligence team.
That worked out surprisingly well initially; the team was closely related to one of my passions, observability, and since Karl had a program management background he made sure I had PM support. That’s where I learned the lesson that if you want executives to care about incidents, you just need to overlay a graph of revenue with your downtime graph. OK, it’s more complicated than that, but that’s the core of it: find the effects your audience cares about and tie them to the problems you care about. Having my own BI team made it easy to get the resources I needed for this.
I also wound up getting to work closely with a bunch of the team from Dyn, which Oracle had recently acquired to provide Oracle Cloud with a DNS offering. Great folks who were based near my youthful New Hampshire home.
Most importantly, while I don’t recall the exact order of operations, I leveraged Karl’s ownership of operational quality at Oracle Cloud into becoming the facilitator of the weekly operational excellence meeting.
That meeting was modeled after the AWS weekly operational excellence meeting. All engineering directors and managers were expected to attend. The meeting opened with a quick round table where teams summarized any high severity incidents that had happened in the last week, followed by spinning a virtual wheel to select a team to dive deeply into their metrics and dashboards, followed by team presentations for any recent high severity incidents that would benefit from wider review of the CAPA.
CAPA: Corrective And Preventative Action. If you’re an Amazonian, think CoE. If you don’t need a special acronym to make you stand out, it’s a postmortem. If your HR department finds “postmortem” to be too depressing, it’s a retrospective.
I loved running that meeting a lot, and I don’t usually like facilitating meetings all that much. I think I had fun because I really pushed the meeting back towards being blameless. Amazon wasn’t good at operational excellence because people yelled at you frequently, it was good at it because standards were high, which is a tricky but important distinction. And one place Amazon failed (at least for me) was that it didn’t provide a ton of incentive for people to help each other be excellent; everyone was running at full speed all the time.
Oracle Cloud had a little more room to breathe, so I had room to shift the culture a little bit. The biggest thing I did was to add a section to the summary of recent incidents. Formerly it’d been a one-way street with teams summarizing their recent issues, often with some embarrassment. I added a stock question: “What can other teams learn from this incident?” You may recall my thoughts about finding common problems.
Now, experiencing an incident was an opportunity to help other teams. Now, everyone had a reason to pay attention – you never know what subtle Java bug is going to trip you up until you trip on it, right? So it’s better if someone warns you in advance. Helping other people is a great way to win trust.
That, plus cutting off anyone who got overly negative during the metrics review or CAPA review, worked well. The metrics review segment became an opportunity to learn from people doing a good job rather than an opportunity to criticize people who were doing poorly. I was slowly, instinctually, creating a web of trust among the engineering leaders in that room. So satisfying.
Unfortunately, the rest of my Oracle life was not going so great. Around this time, someone at the Oracle mothership decided it was time to shut down the earlier iteration of Oracle Cloud. I think this was around the same time Don Johnson, founder of my version of Oracle Cloud, was getting promoted and Thomas Kurian was reading the writing on the wall and (eventually) heading for Google Cloud. This meant that there were a lot of directors at Oracle Cloud 1.0 who needed a place to land.
One of those places was under Karl. A director who had been managing a BI team in that original Oracle Cloud moved over to our team and got my BI team, which was disconcerting. A very short while later, like a month or so, that director quite reasonably found another gig and I got the operational BI team back. A little bit after that I got another BI team, which didn’t really have much to do with operational BI. They were a much more traditional business-focused data team. I really liked their lead but they weren’t a great fit, plus I was gonna have to decide who was going to be the lead once I merged the two BI teams.
It was a lot, and despite the fact that Karl was about to recommend me for promotion – I definitely wanted to get back to director level, for pay and for scope – I decided I was about done trying to balance the semi-useful SRE team (remember them?) plus the operational excellence work I was loving plus the BI quandary I’d inherited. So I started looking yet again, and that’s how I wound up at Zillow. Tricky interview given that Zillow was in the same building as Oracle, the lovely Russell Investments Center. I managed to avoid bumping into anyone during my, uh, vacation day.