
Sign up to save your podcasts
Or


Matty, Trevor and Bridget finally get John Willis on the show, after pestering John since practically the beginning. John has been to more DevOpsDays than anyone else in the room and sits on the core team with Bridget, so the conversation is part history lesson, part organizer's handbook and part year in review. Along the way it gets to burnout, diversity and the code of conduct, and John says burnout beat any technology discussion this year.
John gives the short version. It grew out of BarCamps and open spaces, and Patrick, whom John calls the godfather of DevOps and of DevOpsDays, ran the first one in Ghent in 2009 out of frustration at not being able to get the organizations approached to address devs and ops working together. John says the only American there was probably John. John had been listening to a podcast about agile operations when John almost had an accident, so taken was John with the idea.
The format is usually two days. Mornings have invited keynotes and talks chosen through a CFP, and afternoons are open spaces where attendees make up their own topics and sit in circles. What John loved from the start was the lack of ego, after an early career in rooms full of people with medals of honor who couldn't have a technical conversation: "it seems like everybody checks their ego at the door." Matty has a version of the podcast-in-the-car story too, listening to DevOps Cafe, yelling at the radio that this could never work at Apartments.com, and then hearing Jez tell the HP LaserJet firmware story.
The first DevOpsDays was in 2009, and there were two big ones in 2010, and the plan was to tolerate only two a year, one in Europe and one in Silicon Valley next to Velocity. John says that in the early years only a quarter to a third of the room were first-timers, and the Silicon Valley event was mostly a reunion of the same people. Now, with 22 events this year, about 75% of any given room is new. That is exciting, John says, though it is a little bittersweet. Silicon Valley decoupled from Velocity this year and had about 80% new people, so "Silicon Valley became like everywhere else," and John missed the usual 20 minutes of talking Deming with Ben Rockwood. The upside is local flavor, like the after-party at Target in Minneapolis, and Bridget's "deep dish DevOps" in Chicago.
Matty compares it to Lindy Exchanges from Matty's swing dancing days, which began as local events and turned into the same event with the same traveling DJs in every town. The point is that a DevOpsDays in Des Moines shouldn't be a carbon copy of Minneapolis. Bridget says that is why the core team insists on local organizers, and mentions a group that wanted to hold a destination DevOpsDays in Las Vegas, which didn't come together.
John's cautionary tale is from a few years back. The Silicon Valley event was going to be at Google, and about a month out Google said it wouldn't provide food, then a couple of weeks before it asked every attendee to sign an NDA. The organizers held a red-alert call and moved to a warehouse at the last minute. The next year they tried a hotel, which was a disaster, with a bill that had Patrick calling to ask what in the hell was going on, and John joking that they were about to be hit with "oxygen charges." Bridget adds that "coffee's like $60 a gallon." John's lesson is that volunteers who don't know hotels shouldn't work with them.
Bridget's partner Joe has worked in hotel event technology for almost 20 years, which is why the second Minneapolis event, at a hotel with about 370 people, went smoothly. The first was at a university and they outgrew it, partly because the breakout rooms were on different floors from the main hall and each other, and "we didn't really think it would be a problem and we were wrong." John says the proximity of the open space rooms can make the difference between good and great, and tells of a New York event where the discussion was half done by the time you walked to the room.
On size, Matty says Chicago has sold out both years, that about 90% of people register in the last two weeks, and that keeping it around 300 means you could talk to everybody and everyone shares the same morning. Bridget doesn't think DevOpsDays has to cap at 300, and says Austin is bigger and it depends on local needs. For a first event, Bridget's advice is to look at your local meetup's attendance, since you might get three to four times as many people. Matty and John both advise following the standard format the first time, which Matty frames as the shu-ha-ri idea: "First it's obey, then detach." Chicago stuck with the template the first year and ended up with what Matty calls "exactly bog-standard DevOpsDays," and John adds that Austin, which has run five, took some heat for going to dual tracks.
Trevor is the one host who has never organized one. Trevor has been to parts of two Chicagos and spoke at the second, nervous about talking open source to Microsoft people, and says the crowd was warm anyway. Trevor had also submitted a talk for the first Chicago, and sat in the next room listening to the organizers decide not to include it. Matty says, "I fought for you," and that it was the top call.
That leads John to Jeffrey Snover, whose talk was rejected at an early Silicon Valley DevOpsDays because it was from Microsoft, even though John and Damon had had Snover on DevOps Cafe and were pushing for the talk. On their pre-call, John and Snover discovered they had both worked with Tivoli. As John tells it, Snover was told not to work on what became PowerShell, took a demotion and did it anyway. Matty offers the opposing thought that Snover said on this show that the fight created Snover's resolve, and John gives the Sherlock Holmes and Moriarty comparison. Now Microsoft presentations and sponsorships are common at DevOpsDays, "as it should be," and Snover has keynoted Amsterdam. Matty says Snover is already lined up for Austin.
Trevor made only Chicago. Bridget spoke in New York, ran Minneapolis, and showed up in Silicon Valley for about an hour at a sponsor booth. Matty spoke at the first DevOpsDays Rockies in Denver, did an Ignite in Minneapolis, helped run Chicago again, and gave a talk and an Ignite in Detroit. John did 10 in 2015, including Paris, New York, Austin and Amsterdam, and says Austin, where the organizers are family, is where John's heart is.
John also loved the first government DevOpsDays in Washington, DC, at the US Patent Office, which Nathan Harvey put together. Nathan championed John for a keynote on the history of DevOps, and most of the talks were from people inside government agencies fighting the same battles. John's favorite of the year was Minneapolis: a day early at the Target DevOps Dojo, a lineup with Mary Poppendieck, and an Ignite by John's 12-year-old son on R and data science. Matty followed that Ignite with slides about a Markov bot built from Pete Cheslock's tweets and wondered what kind of jackass Matty was, though John's kids couldn't stop talking about the bot.
The theme John cares about most started at SCaLE early in the year, when John learned that a young man from the community had died, someone John had gotten to know over a few years. It stayed on John's mind for a week. John wrote a blog post about burnout that ran on the IT Revolution blog and got a big reaction, and was then asked to give a keynote on it in New York. Someone who spoke after John, about their own anxiety, told John the article gave them the courage to do it.
Burnout open spaces followed. In Austin, Josh Corman asked a circle of 40 or 50 people how many had seen a therapist, and about 30% raised their hands. John wondered aloud how many would have done so four or five years earlier, and Josh answered, "how many people 40 minutes ago would have raised their hand." Three Velocities in a row have run a burnout session. John says the keynote ended because of not wanting to capitalize on a tragedy, and that the upside is that "things that were possibly hidden in a closet" are now being discussed. Bridget likes the idea of creating humane spaces, so that we don't forget the humanity of people doing this work.
John's other two themes are complexity, with Cynefin, cybernetics, OODA loops and Jeff Sussna's book Designing Delivery all tying feedback loops to complexity, and diversity. John tells Bridget that the post written after Bridget's first DevOpsDays Silicon Valley changed how John saw things. After a great event, somebody said the ops guys should go to dinner, and Bridget wrote about how bad that felt, and John realized how bad it would be to have been the one who did it. John says the community is doing better than most, "but we still suck." Bridget notes that a devopsdays.org event needs a code of conduct: "we won't merge your PR without a code of conduct."
John relays a conversation with a woman in tech who told John everyone talks about getting more women into STEM but nobody talks about keeping them, and John says a Docker core maintainer was hit with harassment from trolls this year. Matty recalls the Minneapolis open space intro, where John, in a single expletive-laden sentence, told whoever was harassing someone in the community to knock it off, and Jon Cowie's talk the next day ended in a 10-minute rant on why.
John's code of conduct story is from the Ghent reunion in 2014. Someone made a bad remark during an Ignite, several people went to Patrick, and the speaker then asked for three minutes on stage to apologize and did. John says the guy wasn't evil, just not being empathetic, and it self-corrected. Bridget calls that blamelessness in action, and says it is the spirit of DevOpsDays.
Bridget's closing point is that "there isn't going to be a done because this is a journey," and that the community is at least looking at itself. John says Jeff Sussna's line "DevOps is empathy" is one John wishes to have come up with, and that it means being bold enough to say in a room full of more experienced people that you don't think Docker works that way, without hearing shut up, you have one year of experience. John tells enterprise people drowning in buzzwords to come to a DevOpsDays and hear practitioners tell real stories about what works and what doesn't.
Matty points to Cora Hays-Magan, whose first DevOpsDays in Chicago was the first day of Cora's dev boot camp and who wrote about it. Bridget's last word is that you don't have to be John Willis or Bridget Kromhout or Matt Stratton, or even Trevor, to run a DevOpsDays, and if you want to run one in your town, the core team can help.
Trevor hosts a one-on-one interview for the first time, admits to being a little nervous, and talks with Matthew Walter about what it's like to switch operating systems. Matthew is a Linux sysadmin at North American Power, a deregulated energy marketer in Connecticut, and has spent the last six months or so moving, partly, toward Windows. Trevor is going the other direction, with more and more Linux clients showing up in Trevor's Windows world. The show opens with Trevor asking whether learning is fun and exciting, and Matthew answering that a struggle would fit too.
Matthew was told in February that the company was moving a new line of business app, which runs most of its backend, from a Linux LAMP stack with PHP to .NET. Matthew is the only ops person there, and was leery. After looking into it, Matthew found a lot that was interesting and a lot of existing knowledge that could still be used, and already ran Ansible for configuration management on the Linux side. Matthew decided to stay, though it was still scary: "I'm a Linux admin and I'm gonna try and learn Windows. It's not something I ever thought I would do."
Trevor has been working with Matthew for a couple of weeks and says Matthew can testify to how much Trevor has struggled with if statements in Bash. Matthew's summary of the learning is that "it's drinking from the fire hose most days."
Trevor's background is Windows and .NET, where installing something means downloading an .exe or .msi and clicking through a wizard. Matthew explains the Linux alternative: if you know the package name, you install it through your package manager, from a reasonably secure source. Packages are signed, come with install scripts, and carry metadata about when they were updated. For anything not in the main repositories there are other people's repositories, like EPEL on RHEL, which someone maintains "through the fantastic goodness of their hearts."
Trevor says Windows is getting package managers too, with Chocolatey and, since Windows 10, OneGet. Matthew has seen Chocolatey and liked it, since it felt like a Linux package manager and "everything just kind of worked," but had heard the packages aren't signed, so wouldn't run it in production. Trevor thinks the repository is curated but the packages probably aren't signed, and tells a story about installing Notepad++ with Chocolatey. The day Notepad++ changed its package directory structure, Trevor's cookbooks stopped converging because the download no longer existed. Trevor's verdict is that it is in its infancy. Matthew adds that with Linux package managers you can mirror the whole repository, which for Ubuntu 12.04 was something like 80 gigs, to control which updates reach your servers.
Trevor brings up the idea, which came up in an earlier episode with Jessica DeVita and Jeffrey Snover, that in Linux everything is a file while in Windows everything is an API or a registry key, which makes infrastructure as code harder on Windows. Matthew says on Linux it is also that everything has one job, small atomic tools chained together, and that "there's nothing that's hidden away." Trevor says every Windows machine needs its Explorer settings changed so it stops hiding files and extensions, and that on Linux you can change almost anything live, while on Windows you nearly always have to reboot, although that is changing. Trevor has been told Nano still needs the occasional reboot but far fewer than earlier Windows Servers.
Matthew asks whether Windows has a driving philosophy. Trevor's answer is that it leans toward one heavyweight tool that does everything, where Linux prefers the smallest possible tool for a job. Matthew says the old stereotype of the Windows admin as unskilled because all they did was click buttons is changing with PowerShell, DSC and Nano, toward tools that assume you know what you're doing. Trevor's counterexample is SQL Server, whose big multi-step installer makes it "like pulling teeth" to automate compared with adding flags to a Linux package install.
Trevor's struggle with Bash was an if statement, and the fact that spaces mean something drives Trevor crazy. Matthew found PowerShell well documented and consistent, and picked it up quickly once learning the Get-Command cmdlet. Trevor puts the difference this way: "I can express my intent in PowerShell, whereas in Bash, I need to know my intent," which leaves Trevor asking what to Google, and Matthew agrees that it's "half Google-fu."
They both wish the old command line would go away. Matthew says it is there for Microsoft's backward compatibility. Matthew's moment of delight in PowerShell came while standing up a VM in Azure: Matthew piped a Get-AzureVM cmdlet, which returns an object of the VM's attributes, into the next command, and it blew Matthew's mind. Trevor warns that some libraries from big companies send back formatted tables instead of objects, which people can read and computers can't use, and Matthew guesses "It's probably a Linux admin doing that." On the Linux side, getting a nice table out of Bash means long chains of pipes with sed and awk, and Matthew says you're often better off going to Python, which is on nearly every Linux box, or Ruby, or Go. Trevor says that after getting used to PowerShell, moving to Bash feels like losing something.
Neither is confident about line endings. Matthew knows the symptom, a file opened on the other OS looks like one line or has an invisible line ending, and is trying to let Git handle all of it: "it's one of those things that will blow up in your face if you don't think about it correctly." Encodings such as UTF-8 are a different can of worms.
On permissions, Trevor says 777 looked like leet speak the first time Trevor saw someone type it. Matthew finds the Linux model simple and elegant, with user, group and everyone else, each able to read, write or execute, and file ACLs for when you need more. Windows gives you full role-based access control from the outset. Trevor realizes that outside a GUI, Windows file permissions are something Trevor has never managed, and Matthew says that is something they will have to check out.
Windows has Event Viewer, and Linux has log files that are appended to line by line, which is easy to ship on to Logstash or Splunk. Matthew's impression of Windows monitoring was that everything is a product you pay for, with support from a closed source company that isn't there for you all the time. Trevor finds both models frustrating, since with open source you may be better off fixing it yourself. Matthew says when the transition started, people kept asking whether they had a support contract for the open source tools, and the answer was "we don't have a support contract with anyone." Their new Windows sysadmin asked "do we have a Microsoft Service Agreement?" and Matthew didn't know.
On LDAP and Active Directory, Matthew says that in a Linux versus Windows contest, "Active Directory wins hands down": the open source LDAP options mostly bolt on features to emulate it, and Matthew got one working, but it was not a fun experience. Trevor has only set up AD in Trevor's own Azure subscription and jokes about seeing the forest for the trees. Matthew's Linux nodes don't use LDAP at all. Ansible connects with a key and no users log in, which Trevor calls hands-off.
Trevor notes that Boot Camp is not always the easiest way to run Windows on a Mac. Matthew says "without VirtualBox and Test Kitchen, life would be much, much worse." Matthew uses a Windows VM to get a PowerShell terminal without spinning up another box, for querying the domain or listing services. Trevor does the reverse, using VirtualBox with Test Kitchen and SSH for Linux, or an Ubuntu desktop when a GUI is needed, and ends up back in Bash anyway.
Matthew ends on the tools: the fact that Chef, Test Kitchen and similar tools now have much better Windows support was the only reason Matthew thought this move was possible. Matthew is learning concepts that apply to both, configuration management, continuous integration and infrastructure as code, which lets Matthew start from what something should look like and work out how to get there, instead of having no idea where things are headed in a whole different paradigm.
Matty catches up with Eric Sorenson, technical product manager for Puppet and the Puppet platform, about Puppet 4, the application orchestration announced at PuppetConf, and, once the two of them agree to put on their pundit hats, Red Hat's acquisition of Ansible. Matty admits up front to not having used Puppet in about three years, so a fair amount of this is a Chef person asking a Puppet person how things work now.
Eric's first CFEngine deployment was in 1998. Before joining Puppet in 2012, Eric built out a Puppet infrastructure at Apple for MobileMe and the iCloud services, and moved to Portland partly for the cycling. Puppet itself is about 10 years old: Eric went back for a PuppetConf talk to the earliest commit in the Git repository, which turned out to be an import from Subversion, dated 2005.
For listeners who only know it as the thing that configures servers, the pitch is that you describe the desired state of your system in a domain-specific language and Puppet enforces that state on the nodes it runs on, in master-agent mode or standalone. Matty says the point is caring about what you want rather than how the sausage gets made. Eric agrees: "I want to have a delicious bratwurst at the end of it."
Puppet 4 is the first major version since Matty last used it. The headline change is a completely rewritten parser for the Puppet language, which had been available for about a year behind a feature flag. Eric got in the habit of calling it the future parser, and now, as Matty puts it, it's the present parser: "Now it's the present parser and the previous parser is gone." The old one came out of what Luke wrote in a series of hotel rooms in 2006 and 2007, and Eric says it had a lot of emergent behavior, some of which people came to rely on and some of which was just odd.
The new one brings features people had asked about for a long time, like loops and iteration and a type system, where a module can declare what it expects passed in, such as a Boolean, one of three strings or a number. Most of it is opt-in, and Eric says Puppet 3 code is pretty much compatible. Puppet 4 syntax starts out looking like a Nagios configuration file, and you can add conditionals and loops as you need them.
On adoption, Eric points to EvenUp, a customer that went all in on the type system with Justin Lambert of the community driving it, and got a big gain in reliability and cleared out a lot of technical debt in their modules. About 17,000 Puppet 4 installations have checked in, though Eric has no numbers on how many use the new syntax. Puppet Forge's quality score, 1 through 5, includes a check that an uploaded module is compatible with the Puppet 4 parser.
The other big change is the agent package. Following Chef's Omnibus lead, Puppet now ships an all-in-one package with a Ruby interpreter, Facter and OpenSSL, unified between open source and Puppet Enterprise, so Eric calls it "the one package to rule them all." Eric says that gives a consistent experience on older platforms like RHEL 4 and, in the commercial version, Solaris and AIX. Eric was surprised how much AIX is out there, and Puppet goes back to AIX 5.1.
Facter, roughly Puppet's equivalent of Ohai, was rewritten in C++ using Boost, and it is fast and still extensible in Ruby or with structured YAML or JSON. That laid the foundation for a demo on the PuppetConf main stage of a prototype of the Puppet compiler in C++, which Eric says is something like 50 times faster than the Ruby one. Matty points out that catalog compilation is the step that runs on the master, and Eric agrees that in agent-master mode "the bottleneck in Puppet is definitely the catalog compilation." On the server side, Puppet has been moving off Apache and Passenger, which could fall apart at large scale with erratic response times, to a stack written in Clojure running on the JVM through Jetty, with the Puppet masters inside JRuby.
Eric wonders what it would look like if Puppet ran fast and cheap enough to be running all the time, converging across the infrastructure almost as soon as you push a change. Matty says some of the organizations Matty works with would be terrified by that, and want it running once a month, which is "scary as hell."
Eric understands where they're coming from. The CFEngine model was to run continuously and revert manual edits immediately, which we now call configuration drift. Eric thinks the shift has been toward understanding drift rather than auto-reverting it in a bastard-operator-from-hell way, with no-op or why-run modes that show what drifted without fixing it until someone acts. Matty adds the argument for frequency: the longer the interval between runs, the bigger the change when the agent finally makes it, and "more frequent, smaller changes really are safer." How often the agent runs should be a conscious business decision and not a limit of the technology.
The product is called Application Orchestration, though Matty prefers choreography, in the Swan Lake sense, and Eric is a Balanchine fan. Eric's counter to the container hype is "even if you have containers, you still need orchestration." Eric hedges on calling it a game changer, since it may be one of the words that causes cringing, but says it might be true here. It started as a prototype Luke wrote years ago, and this year the team productized it.
The idea is to apply the model-based approach Puppet uses for a single node's resources across the application. You describe the components, such as the database server and the app server, how they communicate, and what data one produces for another to consume, like a database connect string that you don't want hardcoded into the app server's configuration. A second part binds those component roles to nodes, which can be machines, VMs or containers. That binding feeds a version of the compiler that builds an environment graph, and a new deployer service walks it, contacts the machines in order, and if an earlier one fails, reports it and aborts the job. Normally you'd run the deployer from CI or the command line after a new commit is promoted, and the nodes talk back over a WebSocket each of them opens to a central service.
Matty is wary of people who think they have an orchestration problem that is really an architecture problem, and of people who want to use it as a script recorder for the manual runbook, with a sysadmin copying binaries and a DBA running SQL. If Puppet or Chef has told a node to install a package, you don't write a check that it got installed, and orchestration should be the same convergent idea with a bigger graph. Eric says there is a mind shift to go through, just as you wouldn't transcribe a bootstrap script line for line into your config management tool. Eric adds that the orchestrator ties together modules you already have or that are on the Forge.
Eric also tells the story of the keynote demo, done live by a teammate named Ryan. A chunk of the configuration was still commented out from an earlier trial run, so in front of 1,500 people Ryan had to open vi and uncomment it. Eric says Ryan went into full Bill O'Reilly mode, with the profane live-demo catchphrase to match.
They announce they are moving to the pundit part of the show, and Matty promises to cut it if it goes badly. Eric's favorite reaction was that Red Hat probably could have gotten Robyn Bergeron back for less than $150 million, and Matty agrees that is the real reason. On the substance, Eric says the business logic makes sense because there was so much Red Hat DNA in Ansible already, and it fits the Red Hat ecosystem of Python tools. Eric also sees the appeal in the low-friction model: you need Python on the box, SSH as transport and a root key, and you can run a sequence of commands across the fleet, which is like capturing an administrator's SSH steps in a repeatable way.
Matty says the ramp-up on Ansible is faster than Puppet or Chef, though Matty figures that runway might run out quickly, and wonders what happens to Windows support, since Red Hat has never had to do cross-platform work. Eric adds that Red Hat has a vested interest in people not running Windows, because they pay Microsoft for a license instead of buying a RHEL license. Eric says the Puppet integration with Red Hat Satellite 6 isn't changing, and that the Ansible piece is a separate layer of command and control, which Eric thinks was the point of the acquisition.
Matty recalls a note from Luke's keynote that fewer than 15% of enterprises use these tools, which is great for people in the space and also a little terrifying. Eric passes along a description from Eric's old boss, Scott Johnston, now at Docker, of the other 85% as whitespace, and says this is why Eric promotes #HugOps: a small number of tool partisans want a deathmatch, and it isn't like that. Matty adds that when a customer isn't ready, Matty would rather say so than sell them the wrong thing. If they buy it, hate it and implement it badly, they won't buy again next year, so it is bad business, and readiness is harder than the technology.
Working a demo booth at PuppetConf let Eric see things a regular job wouldn't show. One was Gareth Rushgrove's work on a Puppet module for managing Amazon resources. The puppet resource subcommand can print Puppet code describing users or files on a system, and with the module it can do the same for AWS. You can build a VPC and instances in the web console, run puppet resource against the API, and get Puppet code you can check into Git.
The other was a talk by Dan Bode on using Consul from HashiCorp to build health checks into resources. Consul publishes a service to its registry only once the check passes, and takes it out when it starts failing, which gives a reactive picture of what is running on the network within 5 or 10 seconds. Eric ties it back to Matty's earlier point about verification being part of the resource.
Matty closes with two stories about HugOps from DevOpsDays Minneapolis. In one, Eric tweeted that when a bystander asked whether Puppet and Chef were going to fight, Eric and Sascha Bates said no, they were going to hug. In the other, Sascha said "friends are more important than where you work." Eric's own metaphor is that "it's all of us together in this tiny little boat trying to get across a giant sea of stupidity."
Various links referenced in the episode!
Matt and Eric are pretty sure the only reason that Red Hat acquired Ansible was to get Robyn Bergeron back.
Matty sits down without a co-host to talk test-driven infrastructure with Arthur Maltson, a software developer who moved into DevOps full time, and Michael Goetz, who manages the Solutions Engineering Group at Chef and identifies as an old-school release engineer with "a lot of personal angst" about unvalidated changes reaching production. Matty warns that the conversation leans Chef-specific because that's what they all know best, but the ideas apply to whatever configuration management tool you use.
Arthur's answer to why bother is confidence: that when you make a change it will work the way you expect. Michael adds that you need to be clear about what you are testing. In Michael's framing there is the signal in (what you told the thing to do), the signal processing (your configuration management tool) and the signal out (what came out the end). You shouldn't test the tool, since it presumably has its own test suite. You should test the things that you and your coworkers are changing on a system. Matty adds predictability: knowing what a configuration change will do in production, instead of doing exploratory testing there.
Matty describes red-green-refactor and says it has been hard to write all the tests first for infrastructure code because of dependencies between pieces. Arthur gently corrects the definition: you don't write all the tests first, you write one or two and then the implementation. For infrastructure that might mean writing a test that a user and group exist, watching it fail, and writing the code to make it pass. "I'm not very religious," Arthur says. "As long as the tests come shortly after or shortly before the implementation code, then you're golden."
Michael splits it into two cases. For greenfield work, Arthur's approach is right, because "you can't test what you don't know." For existing infrastructure that isn't automated yet, some teams write tests against production, using a working system as the blueprint while they develop the automation, which Michael says has been successful for organizations with legacy systems to migrate.
Matty digs into that second case using Chef's audit mode. Audit controls are checks that say if this is true, the system is compliant, and the Chef client can run with audit disabled (the default), audit enabled alongside convergence, or audit only. Running audit-only across a production fleet tells you which machines aren't compliant and what the impact of fixing them would be. In Matty's hypothetical, rolling out convergence code to 10,000 nodes could break 9,000 of them because of snowflakes nobody knew about. It sounds silly, Matty says, but "you're treating your production as your test environment," except that you aren't testing the change there. The results are driving your code change.
Michael says the same pattern works with ServerSpec and other open source tools, as long as you have a known good system to validate against. Arthur ties it to refactoring: "Touching a legacy system that has no test is terrifying," so you write outside-in integration tests first, and in a sense audit mode is essentially rewriting a manual system as configuration management, which Arthur calls a powerful tool.
Arthur walks through the team's setup. A custom chef generate template gives every new cookbook test stubs and Test Kitchen configuration from the start. They use Test Kitchen with ServerSpec and expect to move to audit mode eventually. Speed matters to Arthur as a developer, with feedback in under a minute or two, and the default Vagrant approach was too slow at destroying and recreating machines, so "Kitchen Docker has saved our bacon a bunch." After that comes branch-based development, code review, automated builds, and usually automatic deploy to production once merged.
Michael says the workflow is close to that, with one provocative habit: Test Kitchen instances aren't destroyed until a clean run feels necessary, which will annoy TDI purists. Michael rebuilds from scratch at judgment points to make sure a rebuild works and that a second run changes nothing. The order is a test, then the code, then the next test. Kitchen also gets used with Docker, EC2, DigitalOcean and other drivers. Michael's advice on CI is that people find it scarier than it is: "If you can do it locally, you automate it with robots and your CI pipeline," and you shouldn't add anything in the pipeline that you didn't do locally.
Matty prefers to call it individual development instead of local development, since the workflow could run on shared VM workstations or a cloud driver rather than a laptop. Matty brings up a ChefConf talk by Sascha Bates, whose approach is to run against the existing machine, run it again, and run it again before destroying the VM, because you want to test idempotency against a machine that already has configuration on it. The pipeline should also repeat individual tests, Matty argues, for two reasons: trust but verify, and because your code may by then be merged with someone else's.
Michael says your test environment should look as much like production as you can afford, so 15 production systems means a 15-system test cluster, though Michael lives "in the land of reality" and tells people to validate small chunks so the local footprint stays manageable. Arthur separates fast unit tests from slower integration and end-to-end tests, and says that in infrastructure you have little to play with beyond a beefier machine or test environment. The team's ELK cookbook is tested locally across multiple Docker containers and takes 20 to 30 minutes to get feedback.
Michael's answer to slow runs is that you don't have to build from scratch every time: bake an image that has the earlier steps done and run the cookbooks against that. Arthur gives an example. Arthur's team is rolling out Sensu, and developers built a Docker image with Sensu already in it, so someone writing checks can spin it up and get on with it.
Matty says testing feels like it slows you down, but in practice you move faster because you're not rolling out a cookbook, breaking something, and scrambling to write remediation. "Making things safer overall makes it faster." Matty adds that writing tests forces you to work out the logic, and brings up README-driven development. With customers on a proof of concept, the routine is to have them write down the desired state first, which is effectively the README, then work out the resources, and the tests come out of that. Matty admits to never having written outlines for papers as a student, but says you can't just fire up default.rb and start hacking on infra code.
Michael says complaints about how long tests take grate on Michael: "how much time do you spend on an incident bridge when the thing is broken?" It's almost certainly longer than writing a proper test would have taken. Arthur asks how much time people spend SSHing into systems after a converge to check it worked, and then doing it for the other 100 systems. "You're paying it forward ahead of time by writing those tests."
Michael sees a people gap and a technology gap. The people gap is being too purist and applying decades of software TDD experience, which Michael notes is itself debated, to people new to testing infrastructure, when they'd be better off learning the lessons themselves. The technology gap is fleet-wide validation: there are tools to validate one system, or the output of a web page, but nothing that lets you spin up an app, a web tier and a database and validate all three at once.
Arthur's biggest gap is multi-node cookbook testing. Arthur's team runs the components talking to each other on localhost inside one Docker container, which doesn't represent the real system, and Arthur hasn't seen much written about spinning up a cluster of machines, testing it thoroughly, and tearing it down. Performance and feedback speed come second.
Matty explains that the interview was recorded about 12 hours before Chef Software released a batch of new products, so Matty and Michael added a segment afterward. It is a product overview from two people who work at Chef, and it is not a neutral survey.
Michael says people usually come to compliance because an audit hurt them, and that compliance work means taking a document and translating it into something like "root user must not be 0" and a check to validate it. Matty describes compliance, security and ops teams each working in their own tools, with the compliance folks' "stack is PDF and Excel." Matty's pitch for Chef Compliance is a common language and continuous audit, and it doesn't require the Chef client to be running on the nodes being tested.
The open source pieces are InSpec, a testing framework influenced by ServerSpec, Kitchen-InSpec, which runs InSpec tests in Test Kitchen and doesn't depend on Busser, and Train, an abstraction for talking to local or remote instances over SSH, WinRM, Docker or Mock. Both like that InSpec tests carry metadata such as severity that you can define yourself, and that you can write custom resources so a compliance officer can say what an SSH config should contain without writing a regular expression. Michael has already translated a CIS benchmark for Red Hat, and says a compliance check is really just a test. Matty walks through how it fits together: scan for compliance, remediate with ChefDK, verify with Kitchen-InSpec, run it through Chef Delivery, deploy with Chef Server, and watch with Chef Analytics. "It's not about the tool, it's about how you're doing work," Matty says.
Matty closes by saying InSpec and Kitchen are not Chef-only and work independently of Chef, and that all of it enhances rather than replaces what they discussed earlier. Michael asks listeners to keep an open mind about what they're testing, when, and why.
Bridget hosts this one solo and talks holiday scaling with two people who have been through a lot of peaks. Rob Cummings has spent 10 years at Nordstrom and, since January, supports the operations teams for nordstrom.com. Matt Curry is Director of Platform Engineering at Allstate, and before that spent 8 years at PayPal, or, in Matt's words, "8 delightful holiday seasons of scaling." The twist is that a holiday spike is a known quantity, and the stories are mostly about what people still get wrong when they know it's coming.
Nordstrom has two big peaks a year, the anniversary sale in July and Cyber Monday, with elevated load through the holidays. Bridget asks whether the prediction is a Magic 8 Ball, and Rob says it's "more magic than I would like to admit to." In practice the forecast is worked out with product management as a percentage above last year plus a safety margin. Their habit had been to start testing for Cyber Monday right after the anniversary sale. Since Cyber Monday is the bigger peak, that left too little time to get the kinks out, so they now test for the next peak of the year from the start.
At PayPal the forecasting was more rigorous. Matt says Cyber Monday and the second Monday in December, which eBay called Green Monday, were the biggest days, with a smaller spike in March when people listed the gifts they didn't want. The capacity team used R and other forecasting algorithms to know weekly volume within about 5%. By March, April and May the number barely moved, and the planning went on from there.
Rob describes two methods. They model the traffic in an internal lab, and the models differ because Cyber Monday is mostly anonymous shoppers while the anniversary sale is registered checkout, which hits different systems. They also do what Rob calls "testing in production, because what could go wrong?" That means ramping production-shaped load against production during a non-peak hour and stopping the moment they hit a breaking point. The gap is register checkout and new account signups, which they can't yet test that way in production.
Matt describes PayPal's Holiday Canary Program. Applications got flagged when response time climbed sharply for small increases in throughput, and the flagged ones got canary tests where the team pulled nodes out of traffic or changed load balancer settings to push more load at one host. The bigger challenge, Matt says, was always the giant shared backend resources.
Rob says everyone at Nordstrom cares about the anniversary sale, but Cyber Monday had a culture of assuming it would be fine because anniversary just went fine, so this year took some flag-waving. Rob also changed how performance testing worked. It used to happen at the very end of a dev sprint, right before release, so every problem turned into a question of whether the perf environment or the new code was at fault, and they shipped anyway to find out. Now performance testing is not a gate. Engineering teams are expected to use the perf lab themselves for risky changes, with a full-on test reserved for big complex features.
Bridget asks how you keep speed from letting an index on a query slip through. Matt says PayPal's database team watched new queries and schema changes in staging and ran SQL explains, and that features almost never went in live. They went in off and were turned on over time. Feature flags "are awesome, and they can be terrible if you don't manage them well," Matt says, having seen a feature start corrupting cookies, where turning it off did not put things back to normal. Matt's view is still that restoring service fast matters most: "failure is always going to happen. You just want to make sure the customer doesn't know it's happening."
Rob adds another trick as more of Nordstrom moves to public cloud: send a portion of traffic to the new infrastructure. They routed 10% of product page traffic to it, and performance and add-to-bag rates were worse than legacy. They scaled back to 1%, the team diagnosed the problem and shipped a fix, and the traffic went back up to 50% with performance better and add-to-bag rates where they wanted them. "Having that variable, that slider is handy."
Bridget points out that this works because they measure business outcomes, not only response times. Rob says that has been a culture change, with heavy investment in real user monitoring so they see what the client sees, and it has found anomalies nothing else would. Rob noticed that performance degrades at night when people go home to slower connections. Matt says PayPal could see measurable differences in conversion and cart abandonment based on client time, and Rob confirms Nordstrom sees a link too, though for anything short of a bad performance issue the effect is small.
Rob's answer on freezes is "it depends." Legacy systems that were not built for continuous delivery still freeze, with extra rigor on any fix that has to ship. The newer systems on public cloud are not frozen.
Matt says PayPal called it a moratorium, and moratorium day was the day all of operations threw a giant party. Matt's reasoning is that "the last release and the first release of the year are always the 2 worst": everyone crams features in before the freeze, and then the pent-up features go out at the start of the year. Bridget says that sounds like an argument for small batches, and Rob says Nordstrom sees the exact same behavior. Matt adds that as a payment processor PayPal owed merchants predictability, since the merchants have the same holiday peak.
Bridget asks about the difference between releasing code and making a breaking change. Matt says a PayPal checkout touched something like 80 services, and you never know how the most minor change will interact in production. Matt's example was an eBay seller who became a buyer and ended up with 10,000 addresses, because every address they had ever shipped to became one of theirs, and the system was deduplicating in memory. It "didn't work very well."
Bridget objects that if devs aren't pushing code, entropy and third parties can still break things. Matt agrees, but says a freeze narrows what you have to look at. Rob says incident data on frozen systems shows fewer breaking incidents, and adds that "it's a lot easier to explain to our business partners when a third-party device fails than when we touch something and it broke."
Matt says PayPal ran all-day bridges on every peak day, with everyone in the command center or NOC and hourly checks that everything was green. Rob says Nordstrom's third parties staff up for the peak days, open proactive incidents and review systems, and Matt says the worst that happened with a merchant was a threat to wire PayPal off their checkout.
Matt also says capacity is a sensitive topic around the holidays, because everyone wants to fix every problem by adding hardware, and that can make things worse and cause cascading failure. Matt is honest about the people side too: it makes the CTO warm and fuzzy to see a room full of people watching monitors. Rob says in a 1,600-person technology org some teams are further along with ChatOps and automated monitoring, and others still need the all-day call.
Incident handling itself doesn't change much. Rob says on-calls sit in a room together so escalation is faster, and teams that are not normally on-call get pulled in and escalate sooner. It is the see-something-say-something mentality of "let's just overreact to everything, at least on these couple days." Rob has been pushing the panic button sooner in recent weeks to keep teams practiced, because they only do this a couple of times a year.
Matt's takeaway is that "heroism isn't scalable," and that if you already have a process for finding your risk, you should assess it all year instead of right before the holidays. PayPal also hit years where a software architecture constraint meant no amount of infrastructure would help even at 20% CPU. So they set the target at double what they expected to hit, to force the engineering teams to fix the architecture.
Rob's horror story is from Rob's first year at Nordstrom, a few months in. The whole site came down for the entire anniversary weekend under load. They had never done performance testing and were still on bare metal, so scaling wasn't really a thing. It initially looked like a denial of service attack, and then, "no, no, it's just our customers trying to buy things." Cloud helps, but Rob says the catch is the services that still live in Nordstrom's data center: "I'll tell you what you can't provision on demand, and that's bandwidth." That means dealing with telcos and, sometimes, construction equipment.
Matt says Allstate's worry is a disaster where people need claims and the systems aren't there, which is much less predictable than a holiday. They aren't in public cloud yet, but they run platform as a service, which makes it easier to move workloads and forces them to be more metrics-driven. The cultural work, like rethinking least-access security when job functions change, is ongoing.
Rob says Nordstrom is investing in continuous delivery, infrastructure as code and lean continuous improvement, with plan-do-check-act cycles, and that the customer mobile teams are their unicorns. To spread it, they needed senior leadership bought in, shared weekly demos, and a dedicated team of practitioners who assess whether a team is ready, run a workshop and measure afterward. On measurement, Matt talks about the quality of tests and code coverage on CI servers and about cycle time, and Rob says cycle time is the big one for Nordstrom, where a VP set a goal of reducing it by 20%.
As they wrap up, Rob says the most exciting part of the move to public cloud is teams taking ownership of their systems, so it's "not an ops problem anymore, it's all our problems." Rob wants what Matt has, which is a platform that abstracts some of that responsibility. Matt says of Rob's public cloud developer empowerment, "I want what he has."
Steven Boyd is a certified ITIL expert and a federal employee at the U.S. Patent and Trademark Office, running the service desk's problem management and major incident processes. Trevor opens the show by noting that "DevOps can be a swear word depending on who you talk to," and Matty comes in with a theory: DevOps is the natural merging of Agile and ITIL. This conversation is mostly Steven correcting what the DevOps crowd thinks ITIL is, and Matty and Steven finding out how much they agree.
The pronunciation debate goes nowhere fast. Steven says the spelled-out ITIL is the common one, but that people in the DoD community say it like the word idle, which Matty had never heard. Trevor has never been exposed to it at all, so Steven gives the short version: decades of best practices for IT service management, originally collected inside a UK government organization to cut costs and improve efficiency. The current 2011 version is a repackaging of those practices. Steven stresses that it is "a descriptive framework, not necessarily a prescriptive."
Matty is ITIL V3 Foundation certified, and came to it backward. When Matty took the workshop, the realization was that it was the thing the team at Bank One had been doing the whole time. They just hadn't known they didn't invent it.
Steven walks through the certification ladder. Foundation gives you the vocabulary and an overview. The intermediate certificates split into a lifecycle track (service strategy, design, transition, operations, and continual service improvement) and a capabilities track. The expert level, Managing Across the Lifecycle, is the one that covers the ability to actually create internal processes based on the framework.
Steven's diagnosis of the ITIL backlash is that when someone says ITIL, they usually mean the change management process inside service transition. Matty adds that for most people in ops, change management is the piece they actually touch, through change requests that get rejected while nobody knows what's going on. Steven's explanation for why it is the only piece people know: "Because that's how it was sold." When an organization implements a piece of it badly and brands the result as ITIL, the whole framework takes the blame.
One listener question asked how ITIL's risk aversion squares with fail fast, fail small. Steven pushes back on the premise. ITIL is about classifying changes and documenting them so risk can be assessed, not avoided. Once an organization understands a risk and accepts it, the change becomes a standard change with standing approval, so it doesn't go to the CAB every time. The documentation also gives you traceability when you have to work out which change led to an incident.
Matty takes that in the automation direction: a standard change could be standard because a set of automated compliance, security and test checks are all green, rather than because someone wrote down which buttons to push. Matty wants humans in the loop only where something needs human judgment. That leads to separating risk from impact. Something likely to break that touches one thousandth of your users for five seconds may be fine, while something very unlikely to break that would "set the whole building on fire" may not be. Steven agrees the tolerance is specific to the organization, but says it only works if change management is wired to incident and problem management. Without that flow you have silos of information, "and now you're not making a real risk assessment."
Trevor asks what a CAB is. Steven says Change Approval Board, then corrects that later in the episode: it's the Change Advisory Board, and the advisory part is the point, since the board is meant to receive information back from operations.
Matty ties this to blameless postmortems: people will make mistakes, so the question is how to improve the system. If a change passed every automated check and still caused a problem, the answer is to fix the system, not to tell Trevor the code was awful.
Steven says the major incident process ended in what the DoD calls an after-action report, which Steven treats as akin to a blameless postmortem and as the trigger into problem management. Steven's problem management process starts with a preliminary analysis to scope the investigation and decide whether there is a viable business reason to pursue it. Steven argues that problem management is not only about prevention. It is also about minimizing and mitigating the impact on users, and most organizations invest in change management instead because vendors tell them that is where their incidents come from. In a fail fast, fail small shop you will have incidents anyway, so the question is how quickly what you see in service operations gets fed back into development and design.
Matty describes personal CAB experience as mostly making sure nobody changed a thing at the same time as someone else. Matty would rather trust an automated before-and-after showing that things weren't broken and still aren't, applied in a repeatable way. Matty quotes Mark Burgess: "every time someone logs interactively into a system, they compromise everybody's understanding of that system."
Matty relays a question from Dustin Collins, who had said a shallow understanding of ITIL is that it helps define roles, and who pointed at the failure mode in cross-functional teams: "if everyone owns it, no one owns it." Matty adds an anecdote Matty calls infamous or apocryphal, about a speaker at Etsy asked how people make sure the follow-ups from a blameless postmortem get done. The answer was that they just do, which Matty says doesn't scale.
Steven says ITIL does address this, through the RACI matrix and through the distinction between roles and functions. People like to sit in one silo, but ITIL says an analyst can also hold a role in the change management process. Steven had seen incident teams believe problem management was somebody else's job and that documentation was for the designers.
When Matty sketches how an incident flows into a problem and then into a change, Steven stops Matty with a clarification: "one type of record cannot turn into another type of record." Incidents can trigger a problem, and resolving a problem may require a change, but the records stay separate and are tied together. Matty describes the bridge Matty's team built at Apartments.com between the Agile backlog tooling and change requests, so releases tied into changes and a problem could land back with a product owner as a defect. Steven says ITIL won't dictate that as long as you know your inputs and outputs, and that defect management is a big part of service transition, because without it you can't trace what you see in production back to something testing appeared to resolve.
Matty raises zealotry: saying there is one ITIL way, and that anything else is doing it wrong, is the one thing that would actually be wrong. Matty also doesn't want people to decide they will never have a change board because they had a bad one three jobs ago.
Steven agrees on the zealotry but disagrees a little on cherry-picking. Steven says you shouldn't subjectively pick which parts of ITIL to implement, and you should deviate from the framework only when you are piloting something or you have a process that is better than the basic approach. "It's more of a guide," Steven says, and as long as you stay consistent and document your records, you can modify it. Matty clarifies that the point was extending it and building bridges, not skipping parts. Steven agrees that ITIL gives you the bare-bones best practice and that a better process only creates more value for users.
The last listener question asks how a CMDB fits when infrastructure is volatile and what is worth documenting. Matty's position: "any CMDB that requires manual updates is about as valuable as the bits that it's written onto." If you treat infrastructure as code with something like Chef or Puppet, the tool can populate the CMDB with accurate information, which Matty prefers to a discovery crawl that takes six days and is stale on arrival.
Matty also warns about capturing too much. At the bank, each VM had a CMDB record tied to its host, and since VMware could move VMs between hosts all day, every migration became a change control process because it touched the CMDB. That made automatic load balancing impossible. Steven's answer from the ITIL side is short: document whatever creates value. For problem management, that means enough data to scope an issue, because if you can't scope it you can't assess its impact and urgency, and so can't prioritize it.
Steven's line from the DevOps DC talk, which Matty calls the pull quote for the episode, is a question: would you accept DevOps in a box? If not, why would you accept ITIL in a box?
Matty wanted to spend the last stretch on incident, problem and root cause practices like ChatOps and blameless postmortems. Steven's answer was "We don't have time for that today." They wrap on the point that both ITIL and DevOps are built around continuous improvement, and Steven calls the two symbiotic.
The topic came from Andy Burgin, who runs the LeedsDevOps meetup in the UK and works as a senior DevOps engineer at Sky Bet, and Dustin Collins, who runs the Boston DevOps meetup and is a developer advocate at Conjur. Nathen Harvey of Chef joins them. Nathen says an internal event or community makes sense for the same reasons colleagues don't join external ones: fear of sharing trade secrets, and the practical problem that meetups are after work, when "what you really want to do when you're done with work is go to family or go to pub." A lunch and learn during office hours avoids that.
For external meetups, Dustin says one-way conversations don't help much with hard things like DevOps, since participation is what surfaces issues you hadn't thought of. Andy adds that many people can't get to big conferences, so something local gives them somewhere to engage. Nathen says organizers and participants both have to make sure newcomers feel welcome, because "you're only a newcomer the first time you go." Andy started LeedsDevOps about two years ago because the city's meetups were language-specific and none spoke to Andy's ops work, apart from talks on tools like Vagrant. Andy had also used open source for years without giving back. Dustin took Boston over from an organizer who quit, partly for selfish reasons: to improve at public speaking and organizing, and to meet people.
Nathen's advice is "to not start a meetup group," since it's hard work, and if you're the right person you'll do it anyway. Otherwise, give a DevOps-flavored talk at the PHP group and see whether a few people light up. Dustin found some meetups hated it, but a few resonating people is enough. If the room says "Dustin, you're not allowed to come back," Nathen says, that's permission to start your own. Matty describes suggesting that a suburban organizer run events under the Chicago DevOps group instead of splintering, since it's a lot of work.
Dustin ran Boston alone for eight months, spending six to eight hours a week, until a regular attendee who's very logistically minded offered to handle venues and sponsors. Andy has run Leeds alone for two years by being "essentially lazy": early on Andy took the first two speakers and the most convenient venue, and now the group's reputation lets Andy pick sponsors, with a backlog of speakers. Andy's early marketing was a web page and a Twitter account, following everyone who followed other user groups in the city, and "being quite British and polite, people would follow me back," which brought 100 relevant followers in five days. Andy sees the role as facilitator: "creating the thing and letting the thing happen."
Nathen remembers the first meetups as throwing a party and forgetting the invitations, and says the first DevOps meetup Nathen ran had no topic at all, only a discussion of what the group should be. Nathen ended meetups with checkouts so that everyone shared something and understood "this is not my meetup, this is our meetup." Matty says of the DevOpsDays Chicago kickoff, about 24 people came to the first organizer meeting and about 10 were in the organizer photo seven months later. Trevor says to manage your expectations of volunteers.
Dustin warns that if you start a DevOps meetup with your ops friends, it's easy to end up with an ops meetup: Dustin got about 24 ops people, "and I mean guys," and a 20-minute sidebar on ZFS versus Btrfs. Recruit a developer as co-organizer, Dustin says, and bring in security and business people. Andy says finding speakers beyond Docker talks is hard, and that Windows speakers took extra effort. Dustin adds it's hard to know what people will like: a Python meetup Dustin attended featured a talk about a machine that pricks your finger to test for the flu, with no Python on screen, and the audience loved it.
Matty admits to the echo chamber of inviting the usual suspects. Nathen says if you bring in an outsider, also have a local speaker, since "we should also be building local celebrities." Dustin says to vary formats: long talks, short talks, open spaces and panels, or "it's that time of the month again." Matty says to shake things up while it's going well, since Matty's Chicago events fill within 48 hours and always with the same people. Dustin has speakers start basic and go deep, so both newcomers and experts stay engaged. Nathen says DevOps topics are wide, which is why open spaces work.
Trevor's Azure meetup drew mostly developers, and adding infrastructure talks shrank it. Matty says you can't control your audience, and jokes about putting the dev talk second. Dustin's controversial tip is to bring in a business person with a thick skin to tell it like it is, which happened once as a heckled stand-up and gave technical people a shared enemy.
To a Twitter question about hosting a meetup at a corporation, Nathen says if you want your employer to host, remind them that they're probably hiring, and a meetup brings people through the door. At Custom Ink Nathen would offer an office tour after each meetup, with the aim of getting people excited about working there without pitching them. Trevor adds that it gives your own team speaking practice. Andy says speaking at other groups promotes your own and is how Andy learned what DevOps really was: Andy had thought it was about Nagios and Chef, "and boy, was I wrong."
Matty announces a podcast recommendation of the episode, and starts with The Food Fight Show, which Matty says was a key learning source when beginning with DevOps. Nathen says the idea of checkouts was borrowed from the Ruby Rogues.
Matt spends the entire episode claiming that Nathen was famous for being on ADO11, when in fact it was ADO14.
Joshua Timberman of Chef, Eric Sorenson of Puppet Labs and Robyn Bergeron of Ansible join Matty and Trevor for a panel on treating infrastructure as code. Joshua has been at Chef about seven years with a systems administration background. Eric started by running an ISP "when 14.4K modems were the new hotness," used CFEngine for a long time, and now does product management at Puppet. Robyn was a sysadmin from 1996 to 2000, then Fedora project leader at Red Hat for two and a half years, and has been at Ansible for a month: "systemd is not my fault."
Matty's working definition is that infrastructure is versioned, modularized and testable in an automated way. Eric says at AtomicCon in Portland, seven of ten talks opened with their own definition, and the most compelling, in Eric's view, was that if a natural disaster destroys your infrastructure, you could rebuild it in a new place "using just the contents of a version control repository." Joshua says the original line was Adam Jacob's. Robyn reads the many definitions as a sign that everyone's is shaped by their experiences and should stay open to evolving. Eric adds that seeing scripts checked in, with commit logs you can read like archaeology to divine intent, can be revolutionary for people who have never had it.
Eric says test-driven configuration management didn't exist when Eric started, and credits the Chef community for pushing it. Trevor says people are amazed to find they can check that "the thing I did actually was the thing I did." Joshua likes unit tests because you're testing inputs across platforms, since real environments mix SmartOS, Linux and Windows, and tests catch regressions when someone breaks the most popular platform. Matty says sysadmins are used to testing but not to automating it. Eric quotes an AtomicCon line that everybody has a test environment, and some people are lucky enough to have it be different from production. Robyn's own is "via Twitter when everything goes down."
A Reddit post Matty flags argues that unit tests alone aren't enough for configuration code. Eric agrees they're necessary but not sufficient, and says the state of the art has moved to integration-style tests that spin up machines and assert how they interact, using tools like ServerSpec. Matty borrows a framing from a colleague: signal in is unit tests, signal processing is whether the tool works (the vendors' job), and signal out is acceptance tests, which are "is what I said what I meant." Matty would have newcomers write the "is nginx installed?" test because it teaches the pattern and guards against regressions, and Joshua has come around to 100% coverage of resources so that an accidental deletion during refactoring is a conscious choice.
Matty tells, going by memory of Paul Reed telling it, of Firefox shipping a build that broke MLB.com on the first day of a World Series, and a regression test that has run in every build since. Eric says acceptance tests are your codebase's scar tissue, and that they can slow releases as they accumulate. Joshua says skipping tests just moves the cost, because "you're gonna have to refactor the entire world anyway."
Matty's advice is to start doing it right from day one: even a pipeline with no tests, so the only way anything reaches a server is through source control, because habits are hard to break. Why bother, Matty asks as devil's advocate, when you could keep playbooks on your laptop? Joshua says that works for one person, but real professionals need to recover when a quick change goes wrong. Eric says it's fine "as long as you don't have any coworkers or any customers," and Robyn says otherwise you're making other people have bad days.
Robyn stresses baby steps, since a run of failures demoralizes people, while small successes add up to minutes, and then to a week not spent patching fires. Matty's version is crawl, walk, run: learn to install Apache before you manage the storage array, and pilot with people open to change the way Agile transformations do. Trevor types "trickle-down DevOps" in the chat.
Trevor finds it's often easier to let the infrastructure tool deploy the code, having spent two weeks fighting Octopus over environments. Matty says the application is just another configuration point, but that if you already have a working Capistrano setup you should keep it, and that aligning the two takes a high level of trust. Robyn says it takes trust, transparency and simplicity, and that a group insisting on a tool nobody else has a month to learn is the start of an unhappy relationship. Trevor says don't force everyone onto two tools at once.
Eric describes high-functioning teams where application developers build versioned native packages, like RPMs, and not an obscure tarball or WAR file, so the application sits alongside the OS packages, and calls that what empathy means in practice. Matty notes that people used the empathy talk at DevOpsDays to say the community was all feelings, and points to a line in the Continuous Delivery book that having a very skilled person do mundane tasks is the surest way to ensure error, short of sleep deprivation or inebriation. Eric jokes about a world where Jordan Sissel never had to write FPM, and adds that the tool came out of a low-trust environment where Eric couldn't ask the team to deliver something installable.
Robyn says people at open source companies see trust and transparency on both sides, in the community and in the workplace. Trevor says marketing colleagues at 10th Magnitude were upset by a tweet at a DevOps event asking why companies and their salespeople were there. Robyn's answer is that it's so they can develop empathy for the people using the tools, or it's "a gigantic feedback loop." Robyn adds, "You're either an open source software company or you're not," and that saying you're all about DevOps but not talking to people leads them to expect a unicorn and get a box of Kleenex.
Asked what's changed since they first treated infrastructure as code, Joshua says "everything": you can't be a sysadmin without understanding how to write some kind of code. Eric says culture, with developers building monitoring in as a first-class thing, and tooling, since there are now test frameworks even for Bash. Robyn says open source used to be "no free hippie code" at work and is now acceptable. Matty says the shift is enterprises like Target and GE sharing their stories, after years of being told "you are the 5th person to come in here today and tell me that I should be like Netflix," and Robyn says companies now let employees share theirs, including the failures, in part to keep them.
The Reddit post referenced in the episode:
Having a difficult time wrapping my head around test driven infrastructure
Trevor and Bridget talk with Lindsay Holmwood and Courtney Nash about how human brains shape the way we build and run systems. Lindsay moved from engineering into management a few years ago and has spent about three years reading everything within reach about how people and groups work and about cognitive biases. Lindsay is three weeks into a role as infrastructure and platforms lead at the Australian government's Digital Transformation Office, whose remit is to build "clean, fast, simple, and humane services." The platform is the easy part, Lindsay says, and the hard part is helping teams learn to get the most out of it. Non-Australians can be hired with visa sponsorship.
Courtney, an editor at O'Reilly and Velocity conference chair, was a cognitive neuroscientist in training, a year from finishing a dissertation on learning and skill acquisition, before leaving to work at Amazon. Velocity struck Courtney as operations and therefore not a fit, until finding people who realized that "when you have to start talking about systems of technology, you have to start talking about systems of people." Courtney and Lindsay met at Mountain West RubyConf in 2013 after Lindsay's talk on escalating complexity, which took 60 or 70 hours of research and was emotionally taxing because "people actually die."
Courtney warns that if someone says their product is deep-learning AI, most of that is "just not true," and that the brain tricks us into thinking we make rational decisions. Courtney explains the Stroop effect: it takes longer to say the word blue when it's displayed in green than in blue, because the brain systems for reading words and reading colors conflict. Priming can produce similar friction, for instance in how quickly you associate a woman with the word physician versus software developer. The interesting part, Courtney says, is becoming aware of where those conflicts sit.
Lindsay has given the DevOps Field Guide to Cognitive Biases talk three times. Lindsay describes Kahneman and Tversky's two systems, a fast one and a slow one, with the brain optimizing for speed over accuracy. The guest notes that Thinking, Fast and Slow has one of the lowest completion rates among Kindle books because it's dense, and admits to stopping after a couple of chapters. Smell, Lindsay says, bypasses the fast system, and Lindsay tells of a New Zealand safety company that has workers mount a handkerchief with a loved one's perfume on their equipment to prime them to slow down. Courtney adds that smell takes a different route in the brain than the other senses. Lindsay's gentler recommendation is Dave McRaney's You Are Not So Smart.
Courtney says the good news is that brains aren't fixed. Courtney had "a wicked reaction time" on the Stroop test only from practicing it to embarrass people. Lindsay says biases come from underlying heuristics and can be countered. Lindsay says that in a blameless postmortem you can appoint a devil's advocate to argue for the old view of blaming a person, which gets people to spend more time justifying what actually happened.
Lindsay's other example is the Dunning-Kruger effect: the less you know about something, the better you think you are at it, and you're also worse at recognizing genuine skill in others. The guest says it's what "basically drives DevOps," with ops thinking developers can't run things and developers sure they could keep it up if only they got in there. Even minimal training in a topic improves your ability to self-rate, and being embedded in another team is enough. Lindsay also claims the effect mostly shows up in the West and doesn't replicate in Southeast Asia.
Courtney says empathy has reached buzzword status, and likes that, and recommends Indi Young's work on it as a skill. Lindsay names the three biases that most affect it, as the guest sees it. The fundamental attribution error is judging someone's action by their character, so the driver who cuts you off is an idiot, and not a person swerving for a possum or a wombat. The better-than-average effect: Lindsay says a study in Australia found 70% of people rate their driving above average. The halo effect is when a great impression of a person or brand colors your reading of everything they do, which is how you get a cult of personality.
Bridget asks about Lindsay's talk on the psychology of alert design. The talk's main takeaway is the normalcy bias: before something bad happens we discount it because it never has, and during and after an event we're slow to recognize it. Lindsay cites research on tornado sirens finding people need confirmation from five to seven sources before acting. Courtney relates it to alert fatigue: the brain filters and finds patterns, so if alerts fire without anything bad happening, it stops reacting. Bridget adds the temptation to raise the threshold on an alert that's always going off.
Lindsay was working on an anomaly detection startup until a month earlier, chaining open source tools together to mimic newer findings on how the brain does pattern recognition, like "a basement full of college students that are just looking at a single graph all the time." Lindsay says statistical techniques built on normal distributions don't fit web operations data, and plans to open source the work. Courtney says most so-called AI is really good pattern recognition, which isn't nothing.
Courtney asks Lindsay about early Velocity rejections. Lindsay says the talks were procedural and purely technical, and got better once the work drew on bigger ideas from management, leadership and psychology books, which Lindsay reads instead of Hacker News. The Field Guide talk is, Lindsay says, "a blatant ripoff" of Sidney Dekker's Field Guide to Understanding Human Error, and it worked.
Courtney's advice, borrowed from Scott Berkun, is to give a talk about something you love, hate or are very good at, and to treat the CFP as the proving ground for the idea. Lindsay tries out 5 or 10 minute pieces at local meetups, taking notes the way comedians do. Bridget says speaking at meetups also gets you video, which program committees will watch for unknown speakers. Courtney agrees big stages aren't where to start: "it's terrifying."
Lindsay recommends Patrick Lencioni's books, McRaney's two books, Kahneman's, and QF32, a Qantas pilot's account of landing an A380 and the team dynamics involved. Courtney recommends Ramez Naam's Nexus trilogy and Neal Stephenson's Interface. Bridget recommends the indie film Advantageous and two Velocity Santa Clara keynotes, by Laura Bell and Astrid Atkinson.
From the publisher's feed