The term of art at Amazon for events like Black Friday and Prime Day is “high velocity event.” My entire time at Amazon was a high intensity event. I learned a lot, I appreciate what I learned, I liked the people, I’m grateful to the company, and I never really wanted to go back. Here are some Amazon stories, which I can’t tell without talking about Disney a bit. Bear with me.
I decided to leave Disney because I was bad at the kind of corporate politics required to succeed at the senior director level and I was starting to feel a bit rusty, technologically. I got promoted to senior director because I did some seriously good work to enable selling off a couple of our games, which story I will tell later; I promptly accidentally sacrificed a lot of good will by engaging in discussion about a lateral move without properly preparing my then boss, the CTO of Disney Interactive. Don’t surprise CTOs, y’all. I felt sufficiently burned that when Industrial Light and Magic reached out to talk about a role with them, I didn’t think I was going to be able to navigate the conversation and I turned them down. Some regrets there.
I also wasn’t a great fit with the rest of my peer team. I liked everyone but I wasn’t an executive, I was a technologist who could speak executive fairly well. I wasn’t the kind of ambitious that built a decades-long career at Disney. Because I didn’t really know how to advocate for what I wanted and what I’d be good at, I was getting more and more legacy tech to manage — which meant I had a pretty big team, over 30 people, but they weren’t doing anything people were really excited about.
Writing this, I realize that my boss did try and give me interesting new technological work, which is how I met the ILM folks. This sort of realization is one reason I’m working on these posts; I spent a lot of time being somewhat sour about my last year or so at Disney, and it’s nice to realize that my boss was trying. But again: I didn’t know how to navigate the environment all that well so I didn’t know how to take advantage of the opportunities he probably expected me to grab onto. Lesson appreciated.
As a result of a) missing the opportunities in front of my face and b) getting a lot of legacy tech moved into my team, I felt like I wasn’t getting to play with new tech. I think Docker was the real trigger point for me, although it might have been something similar: I hadn’t been able to get permission to install it even for testing purposes. I was managing a cutting edge MongoDB group, but their tech wasn’t anything we were going to be using heavily in the future, so at this point it was more legacy code. I decided it was time to think about a change.
Ambar, who had previously talked me into coming to work at AltaVista, was my introduction to Amazon. Her skip level manager, David, had been recruiting me for a couple of years with no success but yeah, I was finally ready. By happy coincidence, front-end search and back-end search were merging plus the team which had formerly managed infrastructure for front-end search had mostly all accepted jobs at the hot new Amazon group, Alexa. This meant that front-end search needed a new team, and back-end search had a model which worked very well for them. So David was in charge of hiring a new front-end search technical operations team, which would report into one of his managers and which would span Seattle, Dublin, and Tokyo.
I was a great fit. I was giving up a good title at a stable company but Amazon operated at an entirely different scale and I knew I’d learn a lot. Jason was going to transfer to my team from the back-end search group, and Ambar would be assigned to help me onboard and handle the immediate needs of Prime Day. I prepped my ass off for the interview — I was really worried about my tech chops — and did well enough to get the job.
My first day was as educational as they come. We were in the middle of a high severity incident when I walked in the door, which gave me a chance to see how people reacted to those without the stress of being the one to fix it. It was one of those “only at hyper scale” problems; I won’t get too deeply into details, but we had a dependency on a library whose authors had never considered the case where there were as many servers in a cluster as we had. The day I arrived was the day we went over the magic number, because that’s what the math said we needed for Prime Day in a month, and boom: front-end search at Amazon went down.
For context, that’s the page you see when you do a search. A very significant percentage of people visiting Amazon hit a search page as their first page, because the links you see when you search for a category on (say) Google are links to Amazon search. It’s extremely bad if those aren’t working.
I got a very good picture of how my new teammates reacted to a crisis: calmly and competently. I found out which senior engineers would get dragged off whatever they were working on and into the middle of the problem. I got to see what a good COE (Correction of Error) report looked like. You can’t really arrange for that sort of thing to happen, and educating a new manager is not worth even five minutes of that sort of downtime, but since then I’ve thought that an outbox sandbox exercise would be a smart part of onboarding engineers.
The other thing that stuck with me the first week was orientation, because they introduced the day by saying “this is the largest company you’ve worked for, and we’re all focused on customer service to a degree you aren’t familiar with.” I’ll grant that this was usually true; coming from Disney, I knew that the House of Mouse was in the same ballpark as Amazon as far as size went and they were better than Amazon at customer focus. Regardless, it’s an impressive boast. I took it pretty seriously.
From there on in my days were a haze of preparing for Prime Day. The process was very organized on a macro level. We had weekly meetings with representatives from every team that needed to scale up, with a very clear questionnaire about our readiness status. Watching the lead Availability PM Moustafa run those meetings was a serious education in how to facilitate an availability-oriented meeting.
With Ambar’s help, we managed to be ready for the big week. I only nearly took down the site once during preparation when I was uncertain if my load testing software configuration option was for total traffic generated or traffic generated per load testing server. Since some problems only show up at hyper scale, we had to load test high velocity traffic against our production servers, which meant that I needed a few dozen servers just to generate traffic. It would have been fairly bad if I’d asked each traffic generator to produce enough traffic to stress the full front-end server fleet. But I took one more look at the code before I pushed the button and caught my mistake.
There was an idiosyncrasy in the way we were generating traffic that created unrealistic stress on an upstream dependency that we didn’t catch for a year or so. I still feel bad about that one, although it had been occurring for a while before I got there.
Prime Day itself was incredibly stressful. I didn’t have my team yet so I had to sit on pretty much every war room call for several days running. Our back-end team in Tokyo filled in for me in the wee hours but I knew I was going to be woken up if something bad happened. I didn’t sleep well. The week went perfectly.
I woke up the next working day and realized I was going to have to do it all over again for Black Friday, which would be larger, in just four months. Also I needed to get started on hiring. Also we were going to be transitioning from a pile of legacy Java code into a modern microservices architecture and I absolutely did not have time to provide proper support for that.
I’m stubborn and I liked the people so I dug in, stressing a lot and giving the problems everything I could. My systems did not go down during any of the high velocity events. The front-end search infrastructure team got built out again with a slightly different charter but still wound up providing me with a fair bit of support. Jason dove into preparing for the new microservices and did a really solid job there. I loved working with all the smart front-end engineers and managers.
Hiring was tough. My recruiter wasn’t providing me with a steady stream of candidates so it took a while before I was able to build out the Seattle team to anything bigger than me and Jason. I don’t recommend being on a two person on-call rotation for a major Amazon service. Tokyo got fleshed out pretty quickly, and I managed to hire a nice guy in Dublin; in both those offices, the back-end search operations team helped support us as possible, which also made a huge difference.
Still, I didn’t manage to complete my Seattle team during the entire year and a half I was there. The lesson I finally learned at Zillow was that you gotta do your own sourcing. Even if your recruiter is great, as my Zillow recruiter was, it helps so much to proactively show them what you want.
I did hire a second engineer in Seattle eventually. It’s killing me that I can’t remember his name, but I remember his corgi Snickers. Corgis are herding dogs, and Snickers was an office dog, and he got used to my schedule pretty quickly — so most mornings he’d meet me at the elevator and herd me to my desk, politely and efficiently. On the rare occasions I had 9 AM meetings he got kind of fussy about it, but corgis are also very small dogs so what was he going to do?
I had some successes in the category I love: process ideas. My favorite was an idea to fix the engineering team’s on-call problem; new engineers who didn’t have ten years of context for the old Java search service were understandably having issues handling escalations. So I proposed that we put our five or six senior guys on the old service on-call rotation and tell them not to do any work on anything else during their on-call week; that gave them free time to catch up on old annoying bugs and more importantly, fix the underlying causes of any incidents. We also framed it as a prestige duty. That worked great.
However, I was stressed enough so that I went into therapy for the first time in my life. Having a friendly ear gave me enough perspective to realize that my values didn’t mesh all that well with Amazon’s leadership principles. In particular, I’m a teamwork kind of a guy and Amazon’s leadership principles — good as they are — do not speak much to teamwork. Sometimes people ask me if they should apply to Amazon; I always tell them that Amazon passes the most important test, which is caring about the principles in a meaningful way, but that you have to ask yourself if those principles are the ones you want to live by.
Probably the most important thing I learned from Amazon is that I never wanted to work anywhere that didn’t take core values seriously ever again. You can put your aphorisms on a wall at HQ all you want; it doesn’t count until they’re part of day to day conversation. At Amazon, while I was there, they were.