It’s a point of pride for me that I oversaw technical operations for a pair of very smooth MMO launches at Turbine. Back in the early 2000s, this was by no means a foregone conclusion; in 2004, World of Warcraft had a very rocky (yet successful) launch, and EverQuest had a bad launch in 1999. I’ve had people compliment me on the Dungeons & Dragons Online and Lord of the Rings Online launches at the Game Developers Conference. One person, anyhow. I even did a talk on why Devops is important for online gaming companies at GDC back in 2012, but I can’t figure out where my slides went so I’ll spare you the somewhat outdated evangelizing.
MMO launches were a critical time in a company’s life. Your company probably spent upwards of a hundred million dollars and several years developing your MMO, and you only had one chance to make a first impression. If your game was buggy, that’s what everyone was going to hear about. If people couldn’t play your game, they weren’t going to subscribe. You would never have a better time to capture your audience. One of the powerful things about subscription-based MMOs was the high barrier to exit — people don’t like to give up their established MMO characters thanks to time investment, so you really need to make the most of the launch publicity.
For a little company like Turbine, a bad launch meant the company might shut down. I was painfully aware of this.
These days it’s more routine to launch well. Online gaming has grown up in a big way; if you’re going to spend several hundred million dollars launching an MMO, you’re going to invest some of that in technical operations. Also, the field is way more seasoned. The original networking code for Turbine’s engine was written by a bunch of really smart Brown University students in a basement, and they were shaped by LAN parties rather than the pain of the wide open Internet. They used UDP (fast, unreliable) rather than TCP (slower, reliable) because latency was everything, but because of the unreliability they layered a retransmission protocol on top of it, recreating some of the wheel. To be fair, this enabled a level of twitch-oriented real time combat that may not have been equalled since.
Here’s a story about getting stuck in prior expectations: back when we were in development on DDO, we kept having weird failures between the game servers and the login servers. Couldn’t figure out why. Happened at the same time every day. Rebooting servers fixed it. Finally someone said something like “hey, what do the game servers do when the network connection to the login servers fails?”
“What do you mean, when the network connection fails? Why would that happen?”
Turns out that at start time, the game servers opened up a persistent connection to the login servers and just… assumed it would always be there. Until we took over operations, the login servers and the game servers were in the same data center. In the new network configuration, we were running a VPN between the Bellevue data center where the login servers lived and the Boston-area data center where the game servers lived. (The login servers were shared between all our games, and when we took over operating Asheron’s Call from Microsoft, Microsoft just shifted everything to the Bellevue data center we inherited.) The VPN got reset every night, dropping established connections. Easy fix; not the kind of mistake you’d make if you were a veteran of the Internet. We’ve all gotten much better since.
The other aspect of launching a product that was wildly different back then: no cloud. AWS wouldn’t launch until after DDO and it was quite a while before MMOs were resilient enough and AWS was reliable enough for AWS to make sense. I used to make engineers turn pale by saying “so how will your game worlds react if there’s a guy in the machine room with a .45 who shoots a random server every 48 hours or so?” Not phrasing I’d use today, admittedly: I’d just point at chaos engineering.
Man, the past was a different time.
So how did I launch those two games successfully? Well, I didn’t — an entire team did. The first step here is to launch a fairly successful game like Asheron’s Call and prove the resilience of your code by running that game for several years. Again, Turbine had a very good engineering team. We also had a great QA team; Jay Piette was my peer who ran QA, also hired for the big launches, and he knew his stuff both as a QA leader and as someone who’d spent years working in QA. He was a QA manager at Kesmai, one of the pioneers of online gaming, and had recently run QA for Sims Online. He also took me to a Red Sox game with seats on top of the Green Monster when I moved from Boston down to Maryland; he’s a lovely guy.
I’ve already talked a bit about my direct team in a previous post. I’m gonna shout out Mike Pacheco again, one of the managers on my team: in a previous life he was an Air Force mechanic and the lessons I learned from him are directly responsible for much of our success. Build checklists. Follow them. Double check everything, even the things you don’t need to double check. Mike was a guy who was used to doing technical work that lives depended on; nothing we were doing was that vital.
The core problem with MMO capacity planning was the launch rush. At the time, MMO player concurrency tended to run between 15-20%: at peak, you could expect 20% or so of the subscribers to be logged into the game. Launch weekend was different: concurrency could easily hit 40%, which is a big part of why other MMOs had experienced problems. I did not want to buy twice as many servers as I expected to need in the steady state, and I needed to have a plan for what I called “catastrophic success,” which is what happened to World of Warcraft a couple of years later. What do you do if you get twice as many subscribers as you planned for, and 40% of them are trying to play your game at once? DDO was going to have a soft launch for pre-orders, which helped, but there was still the potential for serious load problems.
My team and I figured out three backstops. The first one was a really clever queuing system that Jon Charette (I think, correct me if I’m wrong) hacked together in a week using literally nothing but our F5 load balancers. Our login system was HTTP-based, and F5 Big IPs had this TCL-based scripting system called iRules. Jon figured out how to use that to create a very basic waiting room, so if the game servers were full people would get dumped into a queue and would be able to log in as soon as there was room. I’m pretty sure the Web team helped out here.
This was not a full solution, because while putting half your player base into a waiting room is better than crashing, it’s still not the impression you want to leave. I needed a second backstop: a way to get new hardware in place quickly.
The data center side was pretty easy; I’d built an excellent relationship with AT&T, who was a major data center player at the time. We had a ton of space reserved in their Watertown facility, plenty of power on tap, and so on. (Sometime later World of Warcraft took space in that facility and by that time we were all using power hungry blade servers, so things got a bit more constrained, but that wasn’t a problem for DDO’s launch.)
What I needed was a way to get servers without waiting a couple of weeks for delivery plus spending the laborious manual effort to get Windows server and all our software installed. I knew how to cut down on install time for UNIX; there may have been better solutions for Windows then but I didn’t have them in my hip pocket, and we had a bunch of manual steps we’d have to go through for each server. The delivery time was the real problem, though.
I was already thinking about this during the server vendor selection process. As part of my due diligence, I flew down to Austin to visit Dell. They showed me a new program they were very proud of; I can’t remember the name but I can remember how well it solved my problem. Basically, you could hand them a server image and they’d preload it onto your servers. What’s more, since Dell was very used to consumer supply and demand patterns, they were able to ship servers much faster than IBM — my other potential vendor.
As designed, this meant I could get servers in a week or so and minimize the configuration time, which wasn’t bad. I wanted more speed, so I negotiated. I think it was an advantage that online gaming was exciting and new; I was a far cry from their usual corporate customers. I also didn’t need that much capacity, relative to bigger companies. I managed to convince Dell to preload several dozen servers with our disk image, pack them up, and have them sitting in a warehouse ready to ship on a moment’s notice. If we’d had a wildly successful launch, I could have those servers arriving at my data center a couple of days later, and game servers up by the middle of launch week. That would do.
Confession: I should have tested this. I didn’t, and it was one of the few parts of the process we didn’t test.
The third backstop I put in place was prepping our vendors. Every single one of them had support on standby for launch weekend. I didn’t want to be left hanging at 3 AM unable to get in touch with a critical vendor. When possible, I had vendors station support engineers in the Turbine offices to reduce lag time even more. Our sole technical issue with the launch involved our big EMC storage unit which kept crashing every night, having EMC support on tap was critical. Turned out to be an old version of the firmware on the disk controllers (and I never missed adding that to my checklists ever again); with EMC’s help, we figured it out within a day or two.
To summarize:
- Have a great team
- Use proven software
- Make checklists and follow them
- Figure out your contingency plans in advance
- Make sure your vendors are ready to handle your worst case problems
Not rocket science, really. We were successful using this plan for two major launches. If I’m being honest, it also helps that DDO was a commercially disappointing launch to the point where I didn’t need to use the Dell contingency; in fact, I didn’t even need to buy more servers for LOTRO since I had left over capacity. I’m still pretty proud of the technical effort.
One final story about the DDO launch, which is when I learned how incredibly invested online gamers can be.
We’d told the fans that the pre-order soft launch would be on Friday, February 24th but we didn’t specify what time so as to minimize the incoming load. I couldn’t sleep all week, so on February 23rd, I was sitting in the office obsessively reviewing everything one more time, talking to whoever else was around, and reading our forums. I realized that some fans decided that when we said we were launching on Friday, it meant we were turning everything on one minute after midnight. Ha! That’s funny; launching when people were exhausted? Terrible plan. I liked the enthusiasm, though.
I kept reading that thread. It was busy. And one guy posting was local to us. So he decided he was going to — come over and take a look? He drove over to our offices, cruised around the parking lot, and came back to report in to his fellow fans. “I don’t think it’s happening at midnight,” he said. “I’ve just been by the Turbine offices and nobody’s answering the front door and I can’t see any lights. Nobody’s there.”
The Turbine office used to be a warehouse. We just didn’t have many external windows so he couldn’t see the lights that were on in my office or anywhere else. But wow. That was a level of dedication I just wasn’t used to from managing search engine Web servers. Little bit cool, little bit uncomfortable.
I decided and have continued to believe that I liked working in an industry that inspires that kind of focus. I’ve also remained aware that some video game fans don’t have good boundaries. Nothing that’s happened in the subsequent two and a half decades has challenged that awareness (cough, GamerGate, cough). I really love gaming, sometimes despite itself.