I am pretty sure this is the specific story that’s gotten me hired more than once. It’s a good story about a project I’m really proud of and I’ve gotten pretty skilled at telling it. It’s also a great example of how rapidly careers could move in those days. I’m good at what I do but I also got lucky here and there and this story is about me being both competent and lucky.
Let’s set the stage. I was hired at AltaVista by Phil Steffora in October of 1998. This was well before I understood how to evaluate a manager either during the interview process or afterwards, so I went in pretty blind. My friend Ambar recruited me, not for the last time, and her recommendation was plenty for me since she worked for Phil as a senior operations engineer. I quickly settled into a comfortable job managing the DEC Alpha servers running our image search tool. I got to work with lots of smart people and it was cool working on the pre-eminent search engine in the world.
In the fall of 1999, the company decided it was time to completely revamp the architecture of the site. When I arrived, we were a simple two-tier system: a bunch of relatively lightweight Web servers sending queries to extremely beefy AlphaServer 8400s, which we called indexers because they created and held the search index. Given our volume, this architecture was straining at the seams and our engineering team decided it was time to move to a three tier architecture with a layer of cache servers between the front ends and the indexers.
I forget the exact numbers, but it was a typical power-law story: the vast majority of our queries were on a small number of terms. If we could serve those queries out of memory rather than from spinning hard drives, we’d be able to run more cheaply and we’d be much faster. Compaq/DEC had just shipped a new rack-mount AlphaServer model, the GS60, which you could get with as much as 12 gigabytes of RAM, so the time was ripe for implementing the cache.
We also decided it was time to move data centers. We’d been running the whole thing in the first floor of a very expensive building in downtown Palo Alto, 529 Bryant Street. The office where I worked was right upstairs, although our software engineers were up in San Mateo, which will become important in a few paragraphs. The Palo Alto Internet Exchange, also a DEC initiative, was in the basement. There was apparently an opportunity to sell the space – I think PAIX needed more room – so moving data centers was a priority.
So: new data center, new hardware, and new software to take advantage of the new hardware. I think we redesigned the UX, too, because why not? The initial plan of record was to just turn off all of AltaVista for a day while we executed on the move. And that’s something we could have gotten away with at the time.
It just seemed wrong to me. I’ve never liked inefficiency. I also knew that the two primary departments working on all these changes weren’t spending a lot of time talking.
My boss Phil was the Sr. Director of Technical Operations. Barry Rubinson was the VP of Engineering. Barry and his team worked up in San Mateo near the 101/92 intersection; if you know the Bay Area, you know that getting between San Mateo and downtown Palo Alto was a 30 minute drive during the day at the best of times, and could be more than an hour at the worst. This was an era before video calls; collaboration was gated by travel time.
I don’t have any insight into how well Phil and Barry got along as people, but from my vantage point their collaboration style was like two rams butting heads over territory. They’d stake out their positions and argue loudly until an equilibrium point was discovered. That point was usually the right one, so I didn’t worry too much about it.
But now I was worried because I didn’t think anyone had thought hard about the fastest possible migration strategy. I decided I was going to take that on myself. I started with my portion of the task, image search. Nick Whyte was the lead engineer on image search so I started driving up to San Mateo and hanging out at his desk asking annoying questions. “Hey, does the new front end for the image search work with the old image search back end? Oh, OK. What about the other way around – if we deployed the new back end and pointed the old front end servers at it, would that work? Could it be made to work?”
From there I got Nick to introduce me to the right engineers to talk to about the rest of the architecture. “Hey, could the new Web search front end talk to the indexers without a cache layer in between?”
Then I started quizzing the smart guy whose name I’ve forgotten who we contracted to project manage the new data center buildout: “When exactly are the new servers going to be available? Uh huh. And we’re moving a bunch of the indexers… when?”
Then I went back to Ambar and started asking capacity questions. “What happens if we run on half the normal number of indexers for a day… oh wow. OK, could we do it during our low traffic periods?”
I now know I was building a dependency graph. At the time I thought I was just drawing boxes and lines. It took a while. When I was done I asked Phil for some time and opened up that meeting with a confident “So I think we can do this move with only two hours of downtime, not a full day.”
Phil looked at my boxes and lines, asked me a few questions, nodded, and said something like “Great, go for it.” I said “Huh?” He said “You made a plan. It’s a good plan. Drive it, and tell people you have my blessing.”
So I did. It was really stressful. I moved a futon into my office for the last week or so, and I used it a couple of times. (Yeah, back in 1999 individual contributors still sometimes had offices.) And in the end we were down for just under an hour – I remember it as 47 minutes, which might be accurate or it might be a number I fixated on. Either way it sure wasn’t a full day or even the two hours I originally projected and I don’t recall any significant issues as we brought the site back up.
A few days later Phil called me back into his office and told me I’d done a good job. Then he noted that I’d made it clear that we needed someone thinking about that sort of problem on a full time basis, so how would I like to be the team’s integration manager? I’d have a couple of people reporting into me – Steve, and I forget who else. I had no idea what I was getting into so I said sure.
I made all the usual mistakes. In particular, I assumed I could keep on doing my old job managing image search servers and still be a manager. Phil didn’t assign anyone else to do image search stuff, after all. I figured that one out after a month or so and randomly talked someone else into doing it. I suspect I was bad at 1:1s.
A few more weeks passed. The Year 2000 bug wasn’t an issue for us that I’m aware of; I spent the evening at the office but I actually rang in the new year at the Rodin Gates of Hell sculpture in Stanford’s Rodin Sculpture Garden, which was cool. In January 2000, I decided to spend a week in London, which was and is one of my favorite cities in the world.
I reminded Phil about this a couple weeks before my departure date. He looked oddly distressed, and the next day he said “hey, can you call me from the airport?” I said sure, I supposed I could, feeling mildly confused.
Cell phones weren’t omnipresent at the time, even in Silicon Valley. I got to the airport early, found a pay phone, and gave Phil a call.
Turns out he was resigning. That was a surprise. Turns out he was recommending me as his successor. That was a huge surprise. Turns out his boss David Henke had decided to promote me, which wasn’t a surprise for me because I had no idea how these things worked, but in retrospect the idea of making me Director of Technical Operations when I’d been an individual contributor less than three months earlier? Kind of insane.
“OK,” I said. “So how long do I have before you leave to learn everything I need to know?”
“Oh, no,” said Phil. “Today is my last day. That’s why I had you call me from the airport – I didn’t want to risk that it might leak.”
I went back to being surprised, but what was I gonna do? I thanked him, headed off to my gate, and had a lovely week in London without worrying as much as maybe I should have.
In the end my lucky break paid off. David Henke took an immense amount of time to teach a very inexperienced leader how to do his job. He taught me that “what gets measured, gets fixed” and I will never forget his speech about how he used to ship compilers at SGI and you don’t get to ship a compiler with a bug in it, so he knew it was possible to ship clean software and us Web guys should stop being so slack.
I also benefited tremendously from working with David Bills, whose exact title I don’t recall but who taught me how to politely tell vendors they were trying to sell me a bill of goods. (“Help me understand that.”)
In return, I did a good job for AltaVista, shepherding Technical Operations through rapid expansion and AltaVista’s ultimately futile bid to challenge Yahoo!. By the time I moved on, I was good enough to get a director-level role without the fortune of having a risk-tolerant boss.