Arrested DevOps

Arrested DevOps

By Matt Stratton, Trevor Hess, Jessica Kerr, and Bridget KromhoutTechnologyTech News
Download on the App Store

Arrested DevOps episodes

  • Tupperware Party with Jérôme Petazzoni, Mark Heckler, and Jennifer Heckler

    Bridget and Matty record at GOTO Chicago with three guests to talk about containers, with Matty playing the listener who knows very little about them. Mark Heckler is a developer advocate on Bridget's team at Pivotal, a Java and Spring developer who also works with Cloud Foundry. Mark and Jennifer Heckler, Mark's daughter and a programmer analyst at Edward Jones, gave a talk that morning on clouds and containers aimed at developers. Jérôme Petazzoni has been at Docker since before it was Docker and once managed a small team of SREs, before giving up the pager to explain Docker and containers to people. Jérôme says that adds up to six years of Docker experience "even though Docker is only 4 years old." Jennifer is the only person on stage who doesn't work at a vendor.

    From Lightweight VM to Just Processes

    Matty asks what the paradigm shift is, since customers often just want to spin up a container instead of a VM. Jérôme says the lightweight VM idea helps people grasp what a container is, but it should be dropped quickly: "it's just processes," or for the technically inclined, cgroups and namespaces. Containers will be many things to many people, the way virtualization went from stacking machines on a host, to cloud APIs, to disposable environments for CI. Matty asks whether calling a container a lightweight VM is itself a wrong statement as a global claim, and Bridget says it depends on the use case. Jérôme sticks to the point, calling the metaphor "just the tip of the iceberg," and noting there are use cases where it doesn't make sense.

    Mark adds the developer view, which is packaging. A VM is "just like driving a nail with a sledgehammer," while a declaratively configured image is repeatable, lightweight and easy to deploy. Bridget adds that a golden image stamps out your application and also a Heartbleed baked in, and Mark says rebuilding and retesting an entire golden image each time a backing service moves is heavy compared with a new Docker build.

    Runtimes and Choosing One

    Bridget asks for the most incendiary line in the talk, and Mark says nothing was flame-worthy, except that some people will not run on a Docker runtime while others will only run on one. Every new Docker release brings cries of anguish when something breaks, and some people are frustrated by the move-fast approach. Bridget asks Mark to explain what a runtime is. Mark sketches runc for initiating containers from images and containerd as the Docker runtime, with other vendors using other execution engines: Joyent Triton, which Mark describes as running Docker containers on Solaris zones, Cloud Foundry, which builds a container per the Docker image format and still uses runc, and CoreOS Rocket.

    Jérôme says Docker's strategy is to give people options, such as swapping runc for Rocket, or SwarmKit for Kubernetes, and compares the choice to picking a hypervisor. For a first cloud VM, you shouldn't spend an hour deciding between EC2 and DigitalOcean, and after hacking at the "cloud jungle" with a machete you'll know enough to choose. Matty agrees that you can't judge engines by letters you don't understand yet, and that "you'll know you need a scheduler when you need a scheduler."

    Diving In

    Bridget asks Jennifer about the actual adoption at Edward Jones. Jennifer says some teams are using containers, and some even run in production, which Jennifer only learned a few weeks earlier. Jennifer describes the financial advisors as the main moneymaker, and the concerns as security and reliability across six time zones. Containers connect to reliability because a sick container is replaced by a new one. Jennifer describes approaching the topic like standing on a beach and looking at the ocean, trying to find where to enter the water, then downloading VirtualBox, Docker, Kubernetes and PCF Dev, following Docker's docs, and running commands like ps to see the output. Jennifer says "You're not gonna horribly break something." Mark adds that Mac users should use the stable builds of Docker.

    Matty, who uses Docker to test cookbooks because a container is faster than a VM, warns that at some point you have to test on something that looks like production, unless production is a container. Mark agrees that less deviation between dev, test and prod is better. On databases, Jérôme asks "should you be running your database in the first place?" If you have one Postgres server and a manual failover in the middle of the night, it probably shouldn't be in a container. If you spin up thousands of databases, say one per CI test, containers pay off. The one deviation Jérôme encourages between prod and dev is running the database in a container in dev when production uses a third-party service.

    Ops Visibility and Buildpacks

    Matty says ops people see a container and ask what the hell it is. Bridget asks whether it holds an unpatched operating system. Jérôme calls the fear unfounded technically, but says techniques that work with VMs don't map to containers, and uses SSH as the example: you can attach a shell to a running container, so there's no need to cram an SSH server and keys into it. Communities, vendors and bloggers exist to share those tricks.

    Mark raises the difference between a Docker image and a Cloud Foundry buildpack, where the container is built around your application. When something like Heartbleed hits, the ops team can patch the underlying platform and reach into those lower layers, so developers don't need to rebuild, and Mark says "That's good and bad, right?" Matty says Habitat is a similar idea of abstracting a layer away.

    Domain Experts and Guardrails

    Bridget asks what to tell a developer who wants to create their thing and not care about buildpacks. Matty says to get off "this full-stack nonsense," since domain experts need to be domain experts, and that whatever is built, be it a Docker image, a Converge node or a binary, is an artifact that should be tested for compliance on its way through deployability. Jennifer says there are two sides: wanting to customize everything, and not having all the time in the world. Matty adds the example of a developer who reads on Stack Overflow that disabling SELinux helps a Node app and changes the cookbook, and says guardrails and fast feedback catch that before security wags fingers. Jennifer says you have to know what's valuable to your company. For a financial firm, that means regulation and privacy, plus user experience teams that once rejected a product for not being user-friendly to financial advisors. Mark says it's a matter of prioritization, developing your expertise and relying on others, and that containers "don't fix your broken culture."

    Aha Moments

    Matty asks for the surprising moment. Jérôme says there was no single one, but a series of crazy experiments, such as running a container on Linux that held a VM showing a screen of Moby running containers inside containers. Matty compares it to Sean O'Meara's configuration management parlor trick of CFEngine installing Puppet that configures Chef. Mark's moment was a gradual realization, from asking why we need this when we have VMs to spinning up Redis or Mongo in containers and seeing the potential. Jennifer says reading Docker docs wasn't enough until building and deploying a simple application, with the realization that you deploy one application in one container, not onto a massive piece of machinery.

    Bridget asks Jérôme for a two-sentence explanation of Moby, and the answer is one: "Docker Inc. is a company that makes Docker, a product that uses Moby, an open source project." Jérôme credits Laura Frank with the explanation.

    Bridget and Matt chat with Jérôme Petazzoni (Docker), Mark Heckler (Pivotal), and Jennifer Heckler (Edward Jones).

    • Mark & Jennifer's GOTO Chicago talk: Clouds & Containers: Hit the High Points and Give it to Me Straight, What's the Difference & Why Should I Care?

    • Jérôme's GOTO Chicago workshop: Container deployment, scaling, and orchestration with Docker Swarm

    • Community & Event Stuff

      If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf

      Upcoming conferences
      Open CFPs
      • lots of DevOpsDays
      • 0 min
      • Risky Business with Nicole Johnson, Matt Curry, and Anthony Lee

        Bridget and Matty record at GOTO Chicago about risk, security and compliance in a DevOps pipeline. Nicole Johnson, who gave a talk on incorporating compliance and security testing into the release process, works with Matty at Chef, and Matty realizes Matty's own "shifting left securely" talk says nearly the same things. The other guests are Matt Curry, a director of cloud engineering at Allstate who leads the organization taking Allstate into the cloud and building its platform as a service, and Anthony Lee, who says Anthony is patient zero for the digital transformation initiative known as Compose, now 2.5 years in. The cold open is Matty's line: "people lie. Computers don't lie."

        Bringing Audit Along

        Bridget asks how an insurance company's customers and regulators shape this. Matt says the compliance and security teams were the tough part of the continuous integration journey, because "explain your job in an algorithm" puts people on the defensive, as if they're being replaced by a robot or a shell script. Anthony adds that insurance means state-by-state regulation plus PCI and SOX. The team reached out to its internal audit organization early and asked them to look at the work, and the greatest outcome was that the lead auditor eventually worked for Matt as a product manager. Anthony relays that auditors saw every production deploy trace back to a GitHub commit, and the reaction was, "I've never walked out of an audit with a smile on my face."

        Matty says every company has compliance with a lowercase c, meaning the standards important to the organization, and everybody thinks they're special. The "dirty little secret" is that one of the biggest ways Chef gets into companies is through audit and compliance, "because people lie. Computers don't lie." Nicole says compliance teams get scared when you go fast without them and hand over a spreadsheet or PDF, but once they're part of the process, they see the value and can collect data programmatically. Matty adds that nobody is automated out of a job, since the risk officer's big brain still decides what's important and the task is describing it consistently. Matt says engaging audit early made communication bidirectional, educated them on what new tools could do, and let the team understand their incentives, which turn out not to be checking the box. That was a chance to "build a bridge rather than kind of pile another brick on the wall," and Bridget notes that bricks into a bridge instead of a wall sounds suspiciously like DevOps.

        Audit Theater and the Compliance Sine Wave

        Matty draws a sine wave of compliance: a company does its regular business and drifts down, then scrambles before the quarterly audit, the auditors show up and leave, and it drifts down again. Bridget asks whether it's really more of a cliff. Matty says that when compliance is part of the process, you're continuously compliant, and a compliance officer knows an auditor could walk in at any time. Anthony describes the typical response as adding process on top of process: one audit finds a missing document, so a process is added to check for the document, and the next audit finds that the checking process failed. "That's what the system solves for you."

        Shifting Left, Hardening Sprints, and Democratized Compliance

        Matty describes a project that ends with a hardening sprint for security testing, which fails because nobody has looked all along, leaving a choice between delaying or getting an exception. Matty's point is that "the bad guys on the internet don't care that you have a note from your mom that says it's okay you didn't patch Heartbleed." Nobody would accept saving QA for the last sprint, and the closer to a defect's introduction you find it, the cheaper it is to fix. Bridget adds what happens if you find out six weeks later, after eight dependencies rely on the hole. Matty says you have to "democratize your compliance."

        Bridget recalls a line from Nicole's talk that Bridget tweeted: a raise of hands for whose job security and compliance are, with the point that all hands should be up. Nicole says it's not a joke. Anyone who touches a system is responsible for compliance, and before you even get to testing, the systems should be hardened, or you reach production with a gold-standard image and find the hardened images break the app. Nicole says you can't make everything compliant right away, so start with the lowest barrier to entry.

        Make the Right Thing the Easy Thing

        Anthony tells of a security team's preferred scanning tool, which Anthony won't name beyond a company that starts with an I and ends with an M, that the devs tried to get into their pipelines and could not. Once security acknowledged it wouldn't work for agile and went with another tool, about 60 dev teams switched in around two weeks, and every commit now goes through the scan. Bridget cites Andrew Clay Shafer's "make the right thing the easy thing," and Matty recalls the example from the book Switch of a machine redesigned so both hands had to be away from the blade.

        Bridget asks what leadership does when it needs to impose a choice. Matt says "if you have a problem that every developer needs to solve, that should become a platform concern," and imagines a world where you don't opt out of certain parts of the CI pipeline. Artisanship is fine inside constraints, and Matt explains the thinking in systems terms: "committees are not scalable," since teams that haven't moved this way can't get a meeting for months. Matty adds that the list of things needing human intervention is shorter than people think. Writing the standard needs a human brain, while checking a system against it does not, and Matty quotes the Continuous Delivery book saying that asking a highly skilled person to do a boring, monotonous task "introduces more errors than inebriation or sleep deprivation." Nicole adds that the standards still need the right humans, those who interact with the systems, to supply context to committees, since you'll never slap a CIS benchmark document on and be 100% done.

        One Pipeline Shape

        Matty says Allstate's decision that there's one way of doing CI and CD means not 60 different teams, and a feature team's core competency isn't building a pipeline. Bridget calls that resume-driven development, and Matty offers "excitement-driven development." Matt says consistency matters for compliance because it makes deployments auditable, predictable and boring. Matty explains the shape stays the same even when Maven differs from another build, so you can look in one Jenkins log no matter the project. Bridget asks Anthony how to motivate people away from what's on the front page of Hacker News, and Anthony says it's a never-ending fight between good and evil, kept grounded by user focus. A team shipped from start to production without talking to the platform team once, which makes it worth it, and constraining the outcome leaves flexibility over tooling.

        Culture, Tools, and Vendors

        Nicole asks how hard it was to change the culture. Matt credits open-minded partners in security and compliance, and says the process is often based on tools already bought, so it becomes a financial conversation, with a sales rep having promised the tool would solve world hunger. Matty adds that organizations tend to throw good money after bad because a decision was made. Bridget notes vendors outnumber customers on the panel. Anthony says that when choosing between Cloud Foundry and OpenShift, the team minimized vendor interaction to judge how well they could run the software on their own, though some vendors won't give access to a download site until a big check is signed. Anthony doesn't advocate cutting vendors off, and credits a great partnership with one.

        Matty says a good vendor partner wants to understand what you're trying to do, not force it a particular way, and Nicole, who is on the presales side, says it means asking the hard questions, such as why do it like that, and serving as a go-between for teams. Matty tells of Bridget referring a contact, who was having compliance problems, to Nicole about Inspect, which surprised the contact because Bridget works at Pivotal and Inspect is seen as a competitor's product. Bridget says the best thing for that customer is often Cloud Foundry, and Cloud Foundry is not Inspect. Matt adds a test for vendors: ask whether they say "yes, totally, I've done this before," or go find the one person in a 10,000-person company who may have kludged it once.

        Closing Advice

        Anthony says to apply the DevOps habit of small, frequent changes to organizations and people: "Small things, small incremental things." Matt says partnership, trust and credibility are the foundation, and that you may have to invest upfront in establishing them. Nicole says everyone is working toward the same goal, even with different paths to getting compliant or secure. Matty says if competitors can hug at a DevOps conference and "give each other hug ops," teams inside one company can too.

        Bridget and Matt chat with Nicole Johnson (Chef), Matt Curry (Allstate), and Anthony Lee (Allstate).

        • Nicole's GOTO Chicago talk: Automating Security & Compliance (for Fun & Profit)
        • Community & Event Stuff

          If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf

          Upcoming conferences
          • Velocity San Jose - discount code "ADO2017" gives 20% off for Gold, Silver, and Bronze passes.
          • Open CFPs
            • lots of DevOpsDays
            • 0 min
            • devopsdays Toronto 2017 with Roderick Randolph, Arthur Maltson, Aaron Aldrich, and Amy Mansell

              Bridget records live at devopsdays Toronto with local organizer Amy Mansell, who is on a podcast for the first time, and three speakers. Aaron Aldrich, a DevOps consultant and community builder at Cage Data in Connecticut, spoke on managing fires and is also organizing devopsdays Hartford, so Aaron is "secretly here for reconnaissance and stealing ideas from other organizers." Arthur Maltson and Roderick Randolph, both at Capital One, co-gave a talk on deep work and structuring a DevOps team. Roderick leads the DevOps practice for Capital One Canada, in what Roderick calls a software studio in North York. The recording ended early, because the batteries in the Tascam recorder ran out about ten minutes before the end.

              Organizing a First devopsdays

              Bridget notes the overlap between ops people and conference organizers, and asks Amy how the organizing team recruited a newcomer. Amy met Steve Pereira, one of the organizers, at a DevOps Toronto meetup about a year earlier, and asked a million questions about the role before saying yes. The appeal was a grassroots event in many cities, and Amy does community building in the day job. Planning one is "a full-time job in and of itself." Amy's pro tip is to talk a lot on Slack, and Amy admits the team may be annoyed by the number of direct messages.

              Steve was too sick to attend, which became a test of the team. Amy says the organizers kept information in the group email and group chat, even for things needed from one person, and held regular meetings and kept shared files. When Steve said Steve couldn't make it, "there was almost just like a playlist that you could go through and check off." Amy adds that the alumni organizers were welcoming, which made the onboarding feel comfortable.

              Foreground, Background, and Deep Work

              Bridget asks Roderick and Arthur how they balance interrupt-driven work against heads-down work. Roderick says to be unafraid to say you need time to focus, pointing to research that "every time you're interrupted, it takes 25 minutes to get back to what you were trying to do earlier." The model from the talk splits a team in two: a foreground group handles the firefighting while a background group works on preventing fires.

              Bridget wonders what happens when everyone who understands Project X is in the background. Arthur says that if only one person can answer a question, that's the problem to begin with, and the team is still working on spreading the knowledge. Arthur hopes to reach a "pair programming or pair opsing model." Bridget pushes back that deep contemplation and conversation are hard to combine, noting that Pivotal is a big fan of pairing. Arthur agrees that foreground pairs can learn from each other, then separate when someone needs contemplation, though the team isn't there yet.

              On onboarding, Roderick says the team timeboxed the model for a month, and it worked well enough to keep iterating, and that being senior doesn't remove the chance to learn from other groups, which is one reason Capital One sponsors the event. Arthur credits Piyush Chugh, a former colleague now at Capital One Canada, with creating the foreground and background idea, and says Arthur took the credit at first.

              Incident Commanders and Backups

              Bridget asks Aaron what happens if the incident commander is a single point of failure. Aaron says the odds are low while an incident is active, but the best practice is to have someone who can step in, and adds that "best practices is a great word we all like to throw around and then pretend they're not real when we actually put them in place." Aaron says onboarding people to the incident process is about leadership setting the tone: being deliberate, and not letting someone work past capacity when they say they're fine, instead saying "I really need you to go home and get rest." The talk came from an incident that went well, and Aaron wanted to capture why, since the company culture side of incident response seemed more important than individual tooling.

              Bridget says ops culture spends a lot of time on what broke and less on what went well, and asks Amy how events handle that. Amy says that on the day of an event, feedback is nearly instantaneous: are people smiling, talking, on time, did everyone eat lunch on time?

              Keeping Questions in Public

              Arthur says the challenge is redirecting conversations that would happen in private chat into a group chat where others can see them, since another engineer who has seen an answer before can give it when the one expert isn't there. Aaron says the team Slack has shifted from private to public messages, in part by reinforcing "default to the public channel," and by copying and pasting private messages to the public channel. Bridget adds that safety matters: if asking a question gets people laughed at, they'll ask in private. Arthur says senior people can use that privilege to show that even someone senior doesn't know everything.

              Bridget chats with devopsdays Toronto local organizer Amy Mansell & speakers Roderick Randolph, Arthur Maltson, and Aaron Aldrich.

              devopsdays Toronto 2017

              Roderick & Arthur
              • Deep Work: A New Working Model For Ops Teams In A DevOps Environment
              • Aaron
                • Managing Fires: The Role Of Leadership In Crisis
                • Community & Event Stuff

                  If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf

                  Upcoming conferences
                  • Velocity San Jose - discount code "ADO2017" gives 20% off for Gold, Silver, and Bronze passes.
                  • Open CFPs
                    • lots of DevOpsDays
                    • 24 min
                    • When the Levee Breaks with Jeff Smith and Mark Imbriaco

                      Bridget and Matty record at GOTO Chicago about communicating in the middle of an incident, with Jeff Smith, a production operations manager at Centro who previously helped build out the SRE team at Grubhub, and Mark Imbriaco, who has spent about 20 years in ops at GitHub, Heroku and DigitalOcean and joined Pivotal "as of yesterday." Jeff gave a talk in Bridget's DevOps track that walked through a live postmortem of a real incident at one of Jeff's unnamed employers, in a complex microservices architecture. Bridget says Mark, a brand new coworker, got dragged to Chicago because of the postmortems Mark has written for places everyone has used. The cold open is Jeff's line, "Don't worry, it's fine. You're fired."

                      Context Is Not State

                      Jeff's big lesson was about context. Alerting on everything can be information overload, so you have to frame the information to tell a story, because you don't want to assemble breadcrumbs mid-outage. The other trap is silencing all the annoying alerts at once, since the landscape of the incident can change, and "you end up fighting the wrong fire." Bridget notes that Bryan Cantrill's keynote made a similar point about alerts that depend on the very systems that are down. Jeff's example is a billing error that got an Azure account shut down, where the only alert was that the website couldn't be reached.

                      Mark says the context problem gets serious in long outages. The Heroku outage ran 67 hours for the last piece to come back, and most companies never think about what happens past 12 hours, when everybody goes into superhero mode. At Heroku they staffed shifts so that two people always worked the incident, offset by 50%, so that with an eight-hour shift, someone new arrived four hours in and could build up context while the previous person was still there. Jeff says a single person carrying an incident start to finish is probably storing all the context, which is dangerous for the organization and the individual.

                      Bridget asks why an oral handoff matters when documentation exists, and Mark separates "context versus state." State is what's up or down on the monitoring page, but the story that led there comes only from the person who lived it, who was too busy to write anything but notes, "marks on trees as they went through the forest solving the problem." Jeff adds that notes record facts, while context is also the theory being chased, and a test script left running overnight means nothing to whoever reads it. Mark adds "you don't write down the negative results," so the next person may walk the same trail.

                      Managing Up During an Incident

                      Matty asks about individual contributors on a bridge with someone seven layers up asking what's going on. Jeff describes Grubhub's common bridge with business and technical people, where business people need to message customers about orders and keep asking how long the fix will take, and "you have to give a number." Matty's answer is the Scotty principle, and Jeff's number is 10, with no units.

                      Jeff's fix is to separate the troubleshooter from the communicator. A troubleshooter gives periodic updates to one person, who becomes the parrot of that message, instead of repeating it each time someone joins the call. The other tool is a running Google Doc, which Jeff thinks was borrowed from the Google SRE handbook. People thought it was crazy at first, then the chat room link to the doc answered the status questions. Mark says the same idea shows up as an incident command role at Heroku and GitHub: the commander captures context and remembers that Jeff is looking at the thing. And "plot twist, that person shouldn't be the one who's communicating externally to your customers either," because choosing how to word a message about bringing a database back up, without saying "recovering" and making people think of backups, takes its own effort. Jeff adds that writing a 15-minute update takes 10 minutes, and that "jumping in to help isn't always helpful" if you haven't checked in with the incident commander, since you could undercut a debugging theory.

                      Matty asks what a sysadmin can do in an organization that has no incident commander, and suggests offering a struggling peer to act as a firewall for communication, then taking the result to management. Bridget suggests a shared doc built from the chat logs, which Bridget says aren't publishable because of the rabbit holes, and building it during the incident might prompt someone to ask whether they missed something with the database. Jeff points out that the first person to grab an incident is already the communicator, scribe and incident commander, so offering to help means asking "what do you need?"

                      Mark's advice for the person being harassed for status is to say "I don't know. I'll tell you something new in 20 minutes," and to give that update even if nothing has changed. Mark compares it to calling the cable company, where the cadence matters more than the answer. Matty calls acknowledgment huge for reassurance, and Mark adds "Not knowing is worse than getting bad news."

                      When an Outage Becomes a Disaster

                      Jeff asks how you mark the point where you concede you have a disaster rather than an outage, since a 20-minute update promises progress and a disaster declaration lets customers stop checking back every 15 minutes. Mark says "I don't know, but I know it when I see it," and describes stretching the cadence: if nothing new is expected in 20 minutes, say you'll update in an hour, or two, and promise to tell them sooner if you can.

                      Mark describes the Heroku case. In 2011 EBS went dark in US East, with a cascading control plane failure across multiple availability zones, and for the first part you couldn't launch apps or restart idled dynos, which took about eight hours to resolve. The long tail was the 200,000 Postgres databases, where some EBS volumes took up to 67 hours to come back or for AWS to give up. Jeff says that with consumer apps you can time a visitor's lifecycle and declare a disaster once the pizza order isn't coming. Bridget says sometimes you know immediately, and describes the 3:00 AM call from an on-call developer saying the HBase cluster was gone and Amazon said it was terminated. Bridget remembers knowing right away that the startup might be done, and it took days to resolve with broken backups, though the company did not go out of business and was later acquired. Jeff's takeaway is "the cloud is just someone else's computer," which doesn't absolve you of recovering from an Amazon failure. Bridget points listeners to the Who Owns Your Availability episode and tells them that if they haven't checked their backups, "you have Schrödinger's backups."

                      A Plan Before You Need One

                      Mark says that joining a company, one of the first things to look at is how it communicates during problems, because it's all about managing expectations. The plan covers frequency of updates, tone for the audience, like Heroku's technical readers versus Basecamp's small businesses, and canned responses. It should be prescriptive enough that nobody makes value judgments mid-incident, usually less than a page, and "you should be completely transparent," since misleading people and later finding you undersold the problem is the worst outcome.

                      Jeff says the way to test an incident process is to simulate incidents. Jeff has been thinking about a talk on pulling the environment's incident creation into the system, like Chaos Monkey for incident management. They ran Failure Fridays and War Room Wednesdays, and Jeff's example prompt is "I just socked Cassandra in the face." Jeff adds that it's best if someone outside the exercise says what's dying, and that with a War Room Wednesday coming you can pre-generate your templates. Bridget warns listeners to clear it with people before taking production down for fun. Matty says to start on paper: "Make your database cluster a sticky note on the board" and pull it off, as good a test as breaking the database. Jeff describes a tool called NetImpair that adds jitter, injects latency or cuts off a node's communications, which simulates failure from that node's perspective. Matty tells a story of turning off the wrong server at Allstate years earlier, with no way to decode the naming scheme, and waiting in the cafeteria to see if anyone came running.

                      Please Don't Be Me

                      Bridget says the reaction inside a company can't be "am I fired?" Jeff describes the week's AWS cost-savings cleanup, in which everyone agreed to delete 67 terabytes of logs on EBS volumes and shut down instances, which broke things. The engineer who ran it was paranoid about being walked out, and Jeff says you can't assume people know they're in the clear, so you have to reinforce it, before repeating the cold-open joke, "Don't worry, it's fine, you're fired." Matty adds that it has to actually be true first.

                      Matty says even when blamelessness is true, nobody believes it until they break something and don't get fired. Matty also says punishing mistakes doesn't reduce them: "It makes them become subject matter experts in hiding mistakes." Matty recalls John Cowie of Etsy on an earlier show asking "how amazing is it when the only thing that happens when you make a mistake is you learn something new?" and cites Etsy's three-armed sweater. Mark says you'll always have the first reaction of "oh my God, what did I just do," but the healthier one is "Please let that be me," because then you know the problem. Jeff says senior people can give air cover by publicly owning mistakes, and Mark notes that a GitHub co-founder wrote a public post about deleting the production database, thinking it was a test database.

                      Bridget says that even at a place that understood all this, the morning after the HBase loss came with the sinking feeling of being probably fired, and Mark says the healthy fear is whether you let customers down, not whether you will lose your job. Matty adds that someone who doesn't care about mistakes isn't someone to run the systems.

                      Jeff says a related burden is being on the hook for not knowing something about a system with 5 million lines of code, and Matty's example is the CTO who asks why you weren't monitoring for that. Jeff says "You're always fighting yesterday's war." Mark points to research by Richard Cook and David Woods: "The way that complex systems fail is fundamentally unknowable," so "Give yourself a pass to be human." Bridget says Tim Gross once sent an email quoting Etsy's principles of blamelessness after an outage when Tim was a director of operations at Drama Fever.

                      Dialects and Postmortems

                      Jeff describes a people-heavy practice: one scribe passes updates to key contacts across the organization, who translate them into their audience's dialect, so developers get a developer version and business gets a business version, and marketing gets what it will publish.

                      Asked for a final word, Mark gives a formula for public postmortems: apologize and mean it, demonstrate a thorough understanding of what happened, and describe what you're going to do, saying "we think it'll reduce the likelihood of this kind of thing happening again" instead of promising it never will. Mark says that can rebuild confidence to a higher level than before the outage. Jeff says to bring the people you communicate with into the process and ask what they needed that they didn't get, and that there are three axes for action items: reduce severity, reduce likelihood, and get better at detection.

                      Bridget and Matt chat with Jeff Smith (Centro) and Mark Imbriaco (Pivotal).

                      • Who Owns Your Availability?
                      • Wide Columns, Shaggy Yaks: HBase on EMR
                      • Jeff
                        • Troubleshooting Tiered Tragedy: A Peek Into Failure
                        • DevOps in the Windy City
                        • Mark
                          • Disasters!
                          • Community & Event Stuff

                            If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf

                            Upcoming conferences
                            • Velocity San Jose - discount code "ADO2017" gives 20% off for Gold, Silver, and Bronze passes.
                            • Open CFPs
                              • lots of DevOpsDays
                              • 50 min
                              • Let’s do the devops again with Nicole Forsgren & Tim Gross

                                Bridget and Matty record at GOTO Chicago with Tim Gross, a product engineer at Joyent, and Nicole Forsgren, CEO and Chief Scientist at DORA. Tim kicked off the DevOps track the day before with a talk on software-defined culture, and Nicole gave a talk on how metrics provide signposts and goalposts on a journey to awesome. Matty, who missed both talks, plays the listener, and notes that recording six episodes in a day leaves about three months of content. The cold open is Nicole's LISA story: "Who do you think pays your bills?"

                                Measuring the Squishy Stuff

                                Matty says it's easy to think you can't put science behind culture: "You can't put a Nagios monitor on a human." Nicole says people often respond by proposing to pull measures about people out of HR systems, which won't capture what teams mean by culture, which is high trust, information flow and collaboration across silos. The alternative is a proxy, and turnover is Nicole's example of a weak one, since someone may leave for a better culture, a worse one, a partner's job, or $1 million. Chat logs work only if Slack is the only way people talk, and Nicole's example of what no system will catch is a coworker dropping off a Diet Coke at a desk.

                                Nicole's answer is psychometric methods, meaning survey questions, done in a research-based way. The Westrum model is the example, named for the researcher Ron Westrum, whose work shows that in high-risk, high-performance teams, a culture that values information flow, high trust, risk sharing and boundary spanning predicts performance. Nicole says the Westrum typology was rewritten, with Nicole's help, into survey questions, six in the short version and seven in the extended one, and that they are open sourced for teams to use every quarter. In the DORA findings over the previous four years it was one of the highest predictors of delivering software with both speed and stability, and it also predicted profitability, productivity and market share. The 2017 State of DevOps Report was due June 15th.

                                Small Teams, Moving Needles

                                Tim has worked mostly in small and mid-sized organizations and worries that a survey would have no meaningful sample size. Nicole says "you can do it with small teams. It still works." Matty, who did this with Nicole at Chef and with customers, says to scope down to the size of a feature team, and adds that Matty doesn't care what number a team lands on, since there's no magic score. What matters is that the parts important to you move, which means asking regularly, not once for a pass or fail.

                                Four Principles of Software-Defined Culture

                                Bridget asks Tim to walk through the four areas from the talk, and Tim says the four are reliability, operability, observability and responsibility, and Tim takes them out of order. Tim jokes that they're a map, not an array, and Matty adds that it will be unsorted every time.

                                Reliability

                                Tim says unreliable software has knock-on effects on the organization. People up all night because of bad on-call burn out and fight with each other, and chasing the shiny normalizes risky decision-making. Nicole says teams need to understand that taking risks is a safe bet and that risks will be shared. Matty cites Charity Majors, who would say that if you haven't broken production, you're not trying.

                                Nicole says the S3 outage was a favorite case, because the published postmortem described a fat-finger incident and "nowhere does it say human error," which speaks well of both the culture and the systems. Bridget takes from that that building for reliability implicitly means not blaming people when reliability falls short. Matty agrees that "you can't work around human error," since people are going to make mistakes and that doesn't scale. Tim adds that many system measurements are themselves proxies for culture. Uptime, for example, is often a proxy for what your people are doing, and Bridget says it may just reflect whether you patch. Nicole says "anything that's a metric becomes a proxy," representing something in someone else's head, and Nicole recalls that as a hardware performance engineer, response time was "my jam".

                                Operability

                                Tim says operability is about delivering quickly and about keeping an application's behavior understandable and self-contained with the team that owns it. Tim objects to the trend of pushing intelligence out of the application into a third party or a platform, because it creates a cultural imperative that it's fine not to understand these things, and it widens the gap between the platform team and the development team. Matty says black boxes let people treat a problem as someone else's, and supply an excuse: if I don't understand how that works, how could I have done it better? Bridget sums it up as "microservices are a game of point the finger and plausible deniability," and Matty says "All of IT is a game of point the finger." Nicole adds "Wait, you mean containers won't fix my culture? What?"

                                Nicole says metrics shape culture and can help teams communicate across boundaries, but they turn problematic when a team throws a container or an app over the wall with a metric that only makes sense to that team. Bridget asks what makes or breaks operability, and Tim says the measurements have to mean something to the consumers of the system, not serve as cover. A vague service uptime number isn't what a consuming team needs. They need to know whether they're being throttled, or whether clients should refresh service discovery. Nicole says to tie metrics to a line-of-business goal, and Bridget cites James Turnbull's argument in The Art of Monitoring that you are not the consumer of your metrics, which Bridget says ops-focused people forget when they build dashboards around what wakes them up.

                                Outcomes, Not NGINX

                                Matty says that outside the echo chamber the simple point still needs making: "Outcomes are like the only thing that matters," specifically the business outcome, and for a nonprofit, Tim adds, that is a mission. Matty recalls a sysadmin freaking out that SQL Server was using all the memory on the server, when that is what the memory is for. Matty also says to write a Chef test for whether a web server does its job and not whether it installed NGINX, and Bridget says the test should look at ports 80 and 443.

                                Matty retells a story Sasha Bates told on the Ship Show about working for a large retailer right before Christmas, when a product team wanted to push a release and Sasha objected on stability grounds. The boss's answer was "your job is not to keep the website up. Your job is to deliver the features that the company needs." Bridget's reply: "I think the company might need a feature of being up."

                                Nicole tells of chairing the LISA conference in 2014, when Courtney Kistler, who had been leading the Nordstrom transformation, stood in for a closing keynote speaker who had a medical emergency. The ballroom of old-school sysadmins was restless about a talk on business transformation, so Nicole and Tom Limoncelli told them to listen, asking "Who the F do you think you're keeping email servers up for?" After about ten minutes of business translation, Nicole says, the room was into it, and "Courtney won them over hard."

                                Responsibility and People

                                Tim says the fourth principle is about externalities. When talking to people about containers, Tim says they don't care about containers: they have a mission, and people whose lives should be fulfilling. Tim adds that people are not just there to fulfill the mission, and on the CEO's duty to shareholders Tim says "fiduciary duty, which is bullshit, by the way. That's not actually a law."

                                Nicole supplies data: over four years, the DORA research found that employees of high-performing teams are 2.2 times more likely to recommend their organization as a great place to work, and research from Harvard found that employees who recommend their workplace predict higher revenue growth. So "even if you want to be a selfish asshole," making the workplace better increases hiring, retention and revenue. Bridget asks the room who is hiring, and every hand is up. Nicole says hiring costs more than retaining, and "Don't be a jerk." Tim says Tim's gut says the same thing, and it's good to see data.

                                Observability and Debuggability

                                Tim says the observability section moves from traditional monitoring toward tools for exploring systems iteratively and collaboratively, not a lone sysadmin watching a dashboard. The stronger point was debuggability, which Bryan Cantrill discussed in the keynote. When a stack has a black box, whether the operating system, a platform or a web server's event loop, people end up saying "and then magic happened." That leaves software less reliable and is dissatisfying for technical people, Tim says, because "This is all software. It's not magic."

                                Nicole adds that it's not enough to have data: you have to act on it, and the highest paid person in the organization "sucks at this. Use your data." For teams without instrumentation, Nicole says to start by asking people. Can you roll anything out without asking other teams? Are you testing, have you shifted left on security? Bridget adds that if you can't deploy a microservice without deploying three others, you may have built a distributed monolith. Matty says to do it iteratively: one question is better than none, and it's information you didn't have yesterday.

                                Nicole says some things only surveys can give you. Systems can tell you what's in version control but not what isn't: "Only your people can tell you what is not in version control," and what is bypassing your systems. Bridget asks Tim about the automation Tim built to capture the state of an AWS account before a migration, and Tim says you can't capture everything, so you start from the top level. That meant documenting the network first, "because nothing else runs without the network."

                                Nicole cautions that instrumenting the easy thing for a quick win creates a pile of metrics, and once something is measured people start paying attention to it. Matty adds that it creates a culture that cares about CPU. Tim asks whether asking people questions has the same effect, and Nicole says it does, because collecting any metric sends the signal that it matters. The problem comes when it turns into a demand to answer 10 on a 1 to 10 scale or else. Matty says the survey for the car Matty had just bought worked that way, and did the same for a Microsoft TAM years earlier, where a scale of 1 to 10 was really pass or fail.

                                Closing Advice

                                Nicole says to start measuring, do it honestly, and keep measuring periodically: "Even a bad baseline is super powerful." Tim agrees and adds that you will chase what you measure, so measure the right things. Nicole is glad they agreed, and Tim says there was no screaming on this stage at all.

                                Bridget and Matt chat with Nicole Forsgren (DORA) and Tim Gross (Joyent).

                                Nicole
                                • Be Awesome With DevOps (Through Data!)
                                • State of DevOps Report
                                • Westrum model of organizational culture
                                • Tim
                                  • Software-Defined Culture
                                  • ContainerPilot
                                  • Community & Event Stuff

                                    If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf

                                    Upcoming conferences
                                    • Velocity San Jose - discount code "ADO2017" gives 20% off for Gold, Silver, and Bronze passes.
                                    • Open CFPs
                                      • lots of DevOpsDays
                                      • 52 min
                                      • Old Geeks Yell at Cloud with Andrew Clay Shafer & Bryan Cantrill

                                        Bridget and Matty sit down live in Chicago with Andrew Clay Shafer and Bryan Cantrill, right after Bryan's keynote, for a panel Bridget pitches as two people with a lot of perspective on where the industry has been. It goes where the title suggests. The next room over asks them to yell less, and Bryan's reply is "It's in the title, it says yell!" Along the way, "Genghis Khan" becomes a code name for Jeff Bezos and "Serpentor" one for Larry Ellison, which Matty asks everyone to use in all future subtweeting.

                                        The Nineties Sucked, and Then Everything Closed Up

                                        Bridget asks how we got into the pickle we're in. Bryan and Andrew dispute the pickle, but Bryan is clear that "the '90s really sucked": a very proprietary, closed era in which people thought systems were done. Andrew points to the dark ages of the relational database and the Java middleware stack that "totally paused everything for a decade." Bryan adds that Java was not open source, and that Windows was deeply proprietary, "then there's asshole proprietary." Bridget sees real willingness to open source coming out of Microsoft now. Andrew says Microsoft has been forced to, since it lost its monopoly.

                                        A New Proprietary Era

                                        Bridget says the proprietary era is over, and both guests disagree. Andrew says that once you're in the cloud, it doesn't matter how much open source built the stuff at Google or Amazon, because you can't change that code. Bryan says "We are in a new proprietary era," and that it rhymes with the '90s, when everything was going to Microsoft and doing anything else was stupid. Now everything goes to AWS. Andrew says "Nothing's more proprietary than Lambda."

                                        Bryan argues that the great myth Amazon created is that cloud is a terrible business nobody should be in, when cloud computing has very good margins and Amazon's overall margin is very low. In Bryan's telling, AWS is underwriting a war on big box retail, and a lingerie retailer that was a Joyent customer answered Amazon's $9 bra with "no, forget it". Andrew says Amazon runs analytics on what sellers and partners do in its marketplaces and turns around with its own products on both the retail and cloud sides. Andrew calls Jeff Bezos "the Genghis Khan of the internet." Andrew remembers Rackspace saying it was playing a different game than Amazon, after throwing a Hail Mary with OpenStack, and Bryan says "Genghis Khan does not follow you on Twitter. He does not care."

                                        The Serpentor detour begins with Bryan raising Larry Ellison's philanthropy and a longevity institute, and arrives at cord blood. It ends with Bridget remarking that vampires and zombies came up earlier in the day than expected.

                                        Lambda, State, and the Borg Paragraph

                                        Bridget asks what the two see among enterprise customers, given all the hype that Lambda functions will save us. Andrew says nothing ever goes away, and describes sedimentary layers of mainframes and Java middleware. Some places are like opening a time capsule, where "you can kind of tell what year they stopped learning." Bryan agrees that mainframes still run but are not a growth area, since nobody holds a conference of several hundred people on z/OS.

                                        Andrew says people over-rotate on running a function. What Lambda represents is a fabric of event sources, and S3 is one, so "you can't build useful things with functions, stateless functions, until you have these things." Bryan says Lambda is "a needle exchange for AWS services," and Bridget adds DynamoDB as the other example of services built to keep you in. Bryan, from "stateless land," is asked whether there's state in the world, and Andrew says there is, along with all the hard problems. Bridget says state is all the customer data and money people do business because of.

                                        Looking ahead, Andrew says new applications should be much more aware of their own state. Andrew recommends the Borg paper, which has pages on schedulers and then a single paragraph that Andrew considers more impactful than the best scheduling algorithm: every application running in Borg has an HTTP endpoint that broadcasts metrics about its health. Andrew says that if you did just that in your applications, "you'd get 85%, 90% of the benefit of the way Google runs their applications." Bryan calls Prometheus the Google idea taken to the open source world. Andrew adds that frameworks should make doing the right thing the easy thing for the app developer, and that people are already struggling to monitor Lambda infrastructure.

                                        Bryan wonders how much Lambda will be used in anger rather than for prototyping, noting that the people who love utility billing are utilities, since you can't predict costs in a utility model. Bridget describes a Minneapolis meetup talk from SPS Commerce engineers, who used Lambda a year earlier and then built their own Lambda-alike in-house because Lambda billing didn't fit. Andrew says that when people cite cost savings from Lambda, it's because they were running mostly idle compute instances.

                                        Open Source Doesn't Die

                                        Bryan says "the open source business model is the second worst business model on the planet," and the worst is being a proprietary infrastructure software company. The Docker CEO had just been replaced that morning, and Bryan recalls the painful period at Joyent when parts of the stack weren't open source. On OpenStack, Bryan says the problem it was solving was a middle management problem in soon-to-be-dead infrastructure companies, and Andrew disagrees that this was the initial goal. Andrew says it became a weird political marketing exercise with little engineering in the core, and that it pulled attention from other promising projects. Andrew's history begins with Eucalyptus, which Andrew says had the birthright to be the open source cloud and mismanaged its community, so that "OpenStack would never have existed if Eucalyptus didn't mismanage its community in the beginning." Bryan says both made the same mistake by trying to be an open source AWS. Andrew's view is that Amazon's advantage wasn't necessarily software but the socio-technical system that had run a massive distributed system for a decade, and that many of the organizations trying to build clouds couldn't manage a multi-node Rails app.

                                        Bryan argues open source outlasts any company: "because open source software can't die." The system Bryan works on descends from Unix, with parts 40 and 50 years old. Postgres was dead on the operating table for a long time, Bryan says, and was revived when things changed. Bryan insists that Illumos has been independent of its Solaris roots for seven years, and that Solaris "is just a very brief proprietary era in a much longer system." Andrew says the Linux and Solaris kernels differ in the quality of engineering, especially around observability. Bryan likes small communities that share values, and says large ones are a mixed blessing. Node.js is the example: Joyent was the company behind it, and Bryan says Joyent and the V8 team shared values around observability, debuggability and rigor that the broader community did not. The community wanted promises, which Bryan calls a bad model "in the operability of software years down the line." Bridget sums it up: day one is very short and day two is forever.

                                        Doing It Right Now

                                        Bryan says doing things properly initially makes you faster than the limit, and that executive leadership has to understand it. Bridget pushes back that doing it right the first time is seductive and impossible, since nobody has perfect future knowledge. Andrew says "there's nothing more expensive than building the wrong thing," and that sometimes testing a hypothesis with a hack that lives forever is cheaper. Bryan says the complaint isn't about prototypes but about the corners you know you're cutting, like being in someone's code that has never been executed and thinking it would have taken an extra 20 minutes. Andrew says every developer on every keyboard is choosing "between doing things right and doing things right now," every second of every day. Bryan passes on advice from a senior engineer that every line of code is a business decision, so an organization has to value not cutting the corner and "be a craftsperson." Andrew adds that values are about who gets rewarded, and that on a resume with stints of 18 months, 18 months, 18 months, each move often came with a big raise. Matty says the same pressure shows up as analysis paralysis among Matty's customers.

                                        Integrity

                                        The reward point sets Bryan off. Bryan says the leadership principles of Amazon and similar organizations leave out integrity, noting that Amazon has 14 leadership principles and integrity is not among them. Bryan contrasts that with corporate values of a generation ago, when integrity topped the list, and says "we have stopped aspiring." Bryan admires a final email to Sun employees saying that in 30 years there had been no need to hide the newspaper from the children, and sets that against Uber's behavior. Andrew asks whether Uber has been rewarded or punished, and Bryan answers that Uber's valuation is pretend. Andrew pushes back that many developers have mortgages and families and are managed like factory workers from the Industrial Revolution. Bryan says they are not, that "our children don't die," and that anyone using their brain in a cube farm is among the haves. Andrew answers that the choices people make to get through the day under imposed management structures aren't something Andrew will begrudge them.

                                        Closing Advice

                                        Bridget asks for a closing statement in 60 seconds. Andrew says "This might not be the podcast you wanted, but it's the podcast you needed," and that integrity starts with individuals: to build structures that have it, you have to think more globally about politics and the implications of decisions, and elevate yourself inside your organization or by creating new ones. Bryan says we live in a great time, making "castles from thought," and "We are literate in a society that is broadly not literate," so it's incumbent on us to increase literacy. Bryan adds "Anyone who bets against humanity is simply ignorant of history," and tells listeners to find their own motivation and take the luxury of being true to themselves. Andrew's last word is that the panel might not have yelled at the cloud, but "we definitely yelled."

                                        Bridget and Matt chat with Andrew Clay Shafer (Pivotal) and Bryan Cantrill (Joyent).

                                        Community & Event Stuff

                                        If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf

                                        Upcoming conferences
                                        • Velocity San Jose - discount code "ADO2017" gives 20% off for Gold, Silver, and Bronze passes.
                                        • Open CFPs
                                          • lots of DevOpsDays
                                          • 53 min
                                          • Enterprises with Bryan Liles

                                            Bridget and Matty talk about DevOps in a large enterprise with Bryan Liles, a director of engineering at Capital One who started as a sysadmin over 20 years ago and has done security, networking and development. Bryan says up front that this isn't Capital One's position, only what Bryan thinks while working there. Bridget met Bryan when Bryan was at DigitalOcean, and Bryan says the appeal wasn't the hot startup but getting back to the basics where Bryan started, at an ISP and a hosting provider, before wanting to do something bigger than Bryan could do alone, which a bank offers. Bridget also apologizes in the intro for audio problems in the recording. The cold open is Bryan's line that "When you have more than 2 lawyers, everything becomes harder."

                                            Start With Why

                                            Matty says that a couple of years ago the statement was that DevOps won't work at big corp, and now the question is how to do it at big corp, which is a fundamental shift. Bryan says how is easy, since you can read a book, and what people miss when they look at Google or Dropbox SREs is why: why they needed that solution and what issues you have. Without understanding why, you never know what success looks like. Even at Capital One, Bryan says, the DevOps teams bring in CI/CD, which is a tiny smattering of what DevOps is.

                                            Matty describes a CIO buying the big-picture message and never getting it to the boots on the ground, who conclude a DevOps initiative means installing Puppet. Matty recalls Jez Humble saying Jez will have a job for life because nobody gets this, and says leaders are now asking people to come to a town hall to explain the principles, not the product. Bryan says to talk about CI, CD, logging, monitoring and declarative infrastructure as principles first, then bring in tools, and Matty asks how you pick a CI tool if you don't know why you're doing CI. Bryan says a big company will just tell you it uses Jenkins, with 20 or 30 people on it, and the conversation loses whether the tool is helping ship code or just being used. Bryan says developers are often given a problem with constraints and told someone else handles ops.

                                            Strategy, Practitioners, and Constraints

                                            Bryan says at a larger org leaders such as directors, VPs and CXOs should talk strategy and where they need to be, and let practitioners figure out how to get there, with a feedback loop. Because tools like Terraform and GitHub Enterprise cost a lot, higher-level people get involved, and should realize they are there for financial support and direction. Matty adds that outcomes are all that matter, and Bryan adds that they must be repeatable, since if you can do it on Tuesday you'd better be able to on Wednesday.

                                            Bryan's view is that "Constraints are a great thing." People who say their startup's problems would go away with more money are naive, since projects with lots of money aren't guaranteed to succeed, and Google succeeded partly because it was trying to do it with fewer people. At Capital One, moving into AWS, they hit account limits, like how many VPCs can be peered, how many images you can copy across regions at a time, and how many security groups you can have. Bryan's example of copying an AMI between accounts, when you have hundreds of accounts and can copy five at a time, means deploying images alone can take more hours than a week has, which forces creative solutions that wouldn't have appeared with unlimited freedom.

                                            Cloud Custodian and Why Enterprises Matter

                                            Bryan says open source is hard at a bank, with lots of lawyers and regulation, but they created Cloud Custodian, a governance tool they released to the community. It enforces 100% encryption at rest: Bryan booted an instance with something unencrypted and got an email 30 seconds later, and again after trying a second time. It also lets you limit things like which instance types can be booted. Matty says that feedback loop through open source and vendors matters, since compliance ideas on a whiteboard differ from a customer's real compliance issues, and someone will be the first at a given industry's problem for Pivotal or Chef.

                                            Bryan says enterprises are also where the money is for vendors, since companies like HashiCorp can't live on VC money forever and have to work with bigger companies. Matty had written off enterprises as where innovation goes to die, and being on the inside as a vendor changed that: it's harder, but the payoffs are better. Bryan says "we're a big, slow-moving ship, but guess what? When we turn around and we actually point in a direction, you get all the force of that ship." The problem is people: ten people have about 100 conversations, and 10,000 people have far more, and Matty says that means 10,000 conversations where good ideas come from. Bryan's favorite thing about working at a big company is that "we have internal tech conferences that are bigger than some conferences that I've been to," where people can talk about specifics without guarding. Matty says a healthcare customer's conference is bigger than every DevOpsDays combined, and that Matty can't go to the talks.

                                            Enterprise DevOps and the Security Door

                                            Matty says "let's all temporarily mourn the term enterprise DevOps and be glad it died," since it was really DevOps without the culture. What Matty sees now is that the entry point into an organization has shifted from ops or dev to security and compliance, since InfoSec people love it. Bryan says governance and compliance are hard at scale, and Cloud Custodian showed there is a marketplace for it, though compliance isn't Capital One's business. Bryan also notes that being in tech at a financial company means not needing to know how the company makes money, and Matty says at Chase the line of business only became clear on updating a resume: $1.3 trillion a day in wires went through the systems Matty managed.

                                            The Changing Role of Ops

                                            Bridget asks about Bryan's Twitter thread with Alice Goldfuss and Kelsey Hightower on ops. Bryan says there are two Bryans, the worker and the person, and the worker can do things wrong without being a bad person. Likewise, "Bad ops teams are teams that don't want to get better," as distinct from overloaded ones, and they need to go extinct. The advice is that you're responsible for your output, and if you can't effect change, "you shouldn't be in a place where you can't make positive change." Bryan says velocity is increasing, and if ops sits on the sidelines saying it's too hard, someone younger will take the job. Bryan is over 40 and expects to keep doing this because times change, from Sun pizza boxes to other people's hardware in data centers. Bryan says to change the solution: if the dev team throws things over the wall, use the SRE book's production readiness tenets and tell them you have production requirements as well as business requirements. Bryan says it's all about being the adult and owning the idea of change for the positive.

                                            Closing Advice

                                            For anyone at a large enterprise, Bryan says to distill a problem to the basics: you're not curing cancer, so solve one little change, then combine them. Startups and enterprises are similar, with a few more rules because more money is on the line. Bryan adds that a big enterprise gives women and people of color more opportunity to move up, and Bryan sees more Black VPs and women at the SVP level than ever before. Bryan leaves listeners with a question: what did you do yesterday, and how can you make today better, and to do something for someone you hadn't yesterday.

                                            Bridget and Matt chat about devops in a large enterprise with Bryan Liles (Capital One).

                                            Check Outs
                                            Matt:
                                            • GFM supports folded details (like, disclosure triangles)
                                            • Tables Generator
                                            • Hugo plugin for Atom written by me
                                            • hub is a cool tool
                                            • Community & Event Stuff

                                              If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf

                                              Upcoming conferences
                                              • GOTO Chicago - Matt and Bridget hosting an entire day of Arrested DevOps Live! May 1-2 - $75 off with discount code "arresteddevops"
                                              • Velocity San Jose - discount code "ADO2017" gives 20% off for Gold, Silver, and Bronze passes.
                                              • Open CFPs
                                                • lots of DevOpsDays
                                                • Image credit

                                                  50 min
                                                • Startups with Charity Majors & Nicole Forsgren

                                                  Bridget talks startups with two returning guests who both started companies after the live episode at DevOpsDays Minneapolis the previous July. Charity Majors is co-founder and CEO of Honeycomb, which is building an observability tool for when monitoring or APM runs into a wall, after Parse and its acquisition by Facebook. Nicole Forsgren is co-founder and CEO of DORA, with Gene Kim and Jez Humble, and before that was at Chef and in academia. Nicole opens with a warning that if you ask about the research, the talking will start.

                                                  The Crossroads

                                                  Nicole says the July conference found Nicole at a crossroads, having left academia just before tenure and spent about a year and a half at Chef, while DORA had been a side project that was getting big. Other companies were approaching Nicole, who didn't know whether to join a bigger one or run the startup. Bridget told Nicole to talk to Charity, whose advice was to decide what Nicole wanted, and if taking a job, to be transparent about the other project and promise a company only a year. Nicole says being about a year out hadn't been articulated until then. Charity says it was clear Nicole was in love with the thing. Nicole took over DORA, committing to Jez and Gene, who had been trying to convince Nicole for months.

                                                  Charity says that at the time of the last conversation, about six months had been spent heads down after a rough co-founder breakup, without knowing whether what they had was worth doing. Around Minneapolis, they decided it was real, and time to push the baby out of the nest. "The first 2 months that you were showing your product to people, you should be humiliated. You should be embarrassed." Now they have their first paying customers and are working through roughly 800 signups from Charity's Twitter feed. Charity says what's hard is to shut up and listen, and that hearing Bridget pitch Honeycomb to others was painful and one of the most valuable moments of recent months. Nicole says the same about the channel partner program.

                                                  Why a Startup

                                                  Nicole says the startup did the hunting. People kept asking to compare themselves with the industry using the State of DevOps data, and some things can only be measured by people, not systems, so Nicole and colleagues built it. Nicole swore never to do a startup and is completely risk-averse. Charity had always been an implementer, not an ideas person, and had assumed a startup would come at some point since it seemed like a missed opportunity in Silicon Valley. The core truth, Charity confesses, is disappointment with the roles offered coming out of Facebook and a successful startup: employers wanted a proof period as an individual contributor or a manager of three people. After a career of saying titles don't matter, Charity found they do, and "it was rage." Bridget is also angry all the time, and it fuels the work. Charity also notes pedigree counts for a lot in Silicon Valley, and that Charity was never going to be more fundable than right then.

                                                  Charity says "Nobody's a startup person," roughly 100% of people, but once in a while something smacks you, and it isn't always the idea. Sometimes it's the team, and the idea comes later, and everyone retcons it as if they always knew.

                                                  Teams and Balance

                                                  Nicole says the partnership with Jez gelled, and "individuals don't make software, teams do," so sometimes when a person leaves, you take the team. Charity says when your strengths and weaknesses mesh with someone's, it's magic. Charity's co-founder Christine says little but is right every time, so Charity copies Christine on everything. Nicole describes the balance of the three founders between focus and creative openness. Charity adds a rule: "Christine and I never freak out at the same time." They joke about whose turn it is to freak out and who's going to be calm in the next conversation. For the first six months it was mostly Charity, who hadn't expected to be CEO, and then Christine couldn't carry them any longer.

                                                  Hiring for a Startup

                                                  Charity says the risk profile is different: at a small startup each hire is one of two or three people and materially shapes the product, which Charity says is true of everyone they've hired, up to seven or eight. You're building a family, technical skills can be taught, but caring about what you care about is hard to tease out and matters more than anything. Charity has never simply interviewed and hired, and who Charity wants to sit next to eight to ten hours a day matters. A big company needs a filter, Charity says, but Charity criticizes Facebook for saying they couldn't lower the bar while having five women in production engineering out of 350, and admitting they found no correlation between interview scores and success. Bridget says a diverse interview panel prevents unconscious bias. Charity says "Great teams can fight well," like couples, and Nicole adds it's how well you resolve conflict. Charity favors transparency below ten employees, including cap tables, and shielding people only from chaos that would distract them.

                                                  Funding and Staying Lean

                                                  Charity says they took $2 million because it was offered on very good terms, are extremely fiscally conservative, and are racing to profitability. When people celebrate raising $100 million Charity sees failure to execute with what they had, and says giving up equity and control is often the unmentioned part. They are raising enough to get comfortably to break-even, so they can walk away from terms. Nicole says DORA is revenue-funded with no investors, which is the best way if you can do it but limits growth, so Nicole uses contractors for accounting, design and copy editing and is feeling scaling challenges. Bridget mentions Honeycomb becoming an ISV partner with Pivotal, and partnerships as a way not to hire.

                                                  Charity says startups need generalists who are fine changing their minds, and that you don't hire into a role until you've done it enough yourself to know what makes someone successful and have six months of work that justifies it. Charity adds that success shouldn't be measured by headcount but by what you deliver divided by the number of people, and is irritated by VCs and acquirers asking how many bodies Honeycomb has, while turning away world-class engineers.

                                                  Advice for Joining

                                                  Charity says the younger the startup, the more ownership comes with it, and you should look for founders aligned with where you want to be. Charity dislikes founders with a "them and us." People should ask far more questions, since many don't know to ask how many shares are outstanding, and look for ways to enrich the relationship before signing, perhaps with a trial period, or by taking friends who know out for drinks. Nicole's advice is to know who you are. Nicole is a big-company person who likes titles and transparency, and jokes about the nickname Crusher of Dreams, earned by spotting things that won't work and needing to be high enough to say so, with a filter that drops in front of the mouth. Charity adds that in your first 10 years you should push yourself, trying big and small companies, and "Running towards being uncomfortable is, I think, generally good life advice." Nicole says give it six months, and Charity says learning to tell good uncomfortable from bad uncomfortable is part of growing up, and that after staying a year and a day at two startups, Charity would now cut it short in a week.

                                                  Bridget chats about startups with Charity Majors (Honeycomb) and Nicole Forsgren (DORA).

                                                  Check Outs
                                                  Charity:
                                                  • Honeycomb has a public announcement/launch on April 11
                                                  • @goturbine for splitting traffic
                                                  • also check out buoyant.io
                                                  • Nicole:
                                                    • DORA has an ROI white paper coming out soon
                                                    • State of DevOps Report releases June 7
                                                    • Book with Jez and Gene; early release is summer/fall
                                                    • Bridget:
                                                      • Systems We Love - coming to Minneapolis March 16th!
                                                      • Community & Event Stuff

                                                        If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf

                                                        Upcoming conferences
                                                        • GOTO Chicago - Matt and Bridget hosting an entire day of Arrested DevOps Live! May 1-2 - $75 off with discount code "arresteddevops"
                                                        • Velocity San Jose - discount code "ADO2017" gives 20% off for Gold, Silver, and Bronze passes.
                                                        • Open CFPs
                                                          • lots of DevOpsDays
                                                          • 56 min

                                                          About Arrested DevOps

                                                          From the publisher's feed

                                                          Arrested DevOps is the podcast that helps you achieve understanding, develop good practices, and operate your team and organization for maximum DevOps awesomeness.