Arrested DevOps

Arrested DevOps

By Matt Stratton, Trevor Hess, Jessica Kerr, and Bridget KromhoutTechnologyTech News
Download on the App Store

Arrested DevOps episodes

  • All Things Docker

    Bridget talks with two people from Docker in early 2021: Justin Cormack, CTO, calling in from the UK, and Donnie Berkholz, VP Products, based in Minnesota. The conversation comes the day after Docker's second online Community Day, which drew 2,500 people, and covers the public roadmap, how product and technology decisions get made, working fully remote, and DockerCon. The cold open is Justin on the scope of the CNCF: "Sometimes it feels like" it covers everything on the planet.

    Listening to Customer Problems

    Donnie joined Docker about five months earlier and has tried to focus the company on customer problems, since tech companies often build a solution and ship it, and then three-quarters of the time it flops. The pain points Donnie hears most are developer productivity and velocity, and finding things you can trust instead of a random repository with three stars or a Stack Overflow copy-paste. Bridget adds the question of whether to take a dependency on something that hasn't been updated since 2019.

    The Public Roadmap

    Justin says Docker tried to change how it labels readiness. Previously enterprise products were marked experimental and you had to "figure out the magic runes" to turn a feature on, which discouraged trying things. The team is culling experimental flags in favor of notices that something is new, and offers early access to developers who want it. The public roadmap has stages, which Justin walks through. Anyone can add an idea as an issue, and product managers look at new items daily, with a weekly meeting that looks at new items and thumbs-up votes. Items move through investigating, writing the code, almost there (asking early users to try it), developer preview or experimental with a public way to try it, and shipped, which means some form of general availability, with feedback still welcome.

    Donnie adds that "investigating" also asks whether a feature is valuable enough that people would subscribe for it, since Docker has a long history of giving lots away and needs a sustainable business model. As an example, the M1 support item became the most upvoted item on the roadmap of all time within days of Apple's announcement. Donnie says Docker isn't a huge company and has to be careful about big engineering investments without a customer signal, so they hadn't started earlier, since it was unclear when it would matter and what else they could be doing.

    CNCF and Open Source Decisions

    Justin sits on the CNCF Technical Oversight Committee, and says cloud native is everything in modern software delivery. The community is large and open, people usually know what's coming, and there's less of a big-announcement culture and more collaboration. Justin says much of the CNCF is about the production end, but more developer-side projects are arriving, such as a Spotify project (whose name Justin forgot). Justin likes assessing projects, since it gives an excuse to ask users how a project is working for them.

    Donnie says deciding what to open source comes down to what the company wants to accomplish, the maturity of that layer of the stack, and whether it's differentiating or becoming a utility or commodity. Open standards help align vendors around things that are becoming a commodity. Bridget notes that end users love any interoperability, anywhere they can just ship that container.

    Remote-First

    Justin says Docker chose to become fully distributed and dropped offices, going from a San Francisco company to one about half European and half US by the end of 2019. The lease on the Cambridge office expired about six months earlier, and it seemed weird to renew it. Docker used to send every new hire to San Francisco for a week, which had become a relic, and the pandemic made the change happen faster. The main limits on hiring are tax, legal operation in different places, and time zone overlap, which Donnie says means deciding whether to look in geographies that overlap or find people willing to work different shifts. Donnie says the one hour of overlap between Germany and San Francisco means you have to optimize for autonomous or asynchronous work, since "force-feeding" synchronous models doesn't work. Donnie adds that online communities and platforms suited to online events reach people where they are.

    DockerCon

    The DockerCon CFP was closing in a few days. Justin says last year's DockerCon was planned as online before the pandemic and drew about 80,000 people, and it showed how much more accessible conferences are without travel. DockerCon is about 80 percent people who identify strongly as developers, unlike KubeCon, where the strongest group works in infrastructure, and Justin wants stories from people who wouldn't normally be heard. The CFP asked about team collaboration, since helping teammates onboard and sharing images is part of the product, but Justin says to submit if it's interesting for developers.

    Donnie says conferences let people swap stories, and that "nobody wakes up saying like, oh, I want to use this tool today." Donnie cites a stat, unverified, that 50 percent of developer time on cloud-native applications might go to configuring instead of writing new code, and Bridget wonders how much of that is usability and how much is security compliance. Donnie describes Docker's experimental Hub CLI tool, and says the interesting part is watching how people use it in pipelines, such as automating token rotation, making sure images have both M1 and x86 builds, or monitoring subscription seats, because people want a use case and not just a tool.

    • Docker Public Roadmap
    • DockerCon CFP open until March 15th - How to Write a Great Talk Proposal for DockerCon LIVE 2021
    • Docker Career Openings
    • Docker Hub Experimental CLI tool
    • 34 min
    • Learning to Learn and Learning to Teach with Shelby Spees

      Matty talks with Shelby Spees, a developer advocate at Honeycomb and former English teacher who has taught and tutored since high school, about learning how to learn and learning how to teach. The topic started when Shelby tweeted about a talk idea and Matty offered to record an episode. The cold open is Shelby: "I needed the 'So, what?' on Kubernetes."

      Growing People Who Live in Production

      Shelby says the idea grew out of a Twitter discussion the year before about what makes an SRE: is there a junior SRE, do you need an Ops background. Shelby wants more people living in production and sees observability as helping. Shelby is wary of being prescriptive, noting that many senior people cut their teeth in server rooms while newcomers have never seen a server rack, and Shelby is a cloud-native engineer who has never worked on-prem. Shelby's own lesson came from owning a tool without releasing a new version, since a year into a career Shelby didn't know how to release a Python package: pretty maintainable code is meaningless if it doesn't reach users. The question is how to teach deployment, maintenance and long-term software life cycle to bootcamp graduates and CS students.

      There's No Linear Path

      Matty says experience is the best teacher but "you can't do unless you have experience," and that it's hard for experienced people to shift their frame. Matty recalls a tech screen with softball opening questions that virtualization-era candidates, who never dealt with hardware, couldn't answer, and trying to learn an API at PagerDuty whose documentation was written by developers for developers, though many consumers were ops people. Matty says Gatsby is built for React developers, which is why it's frustrating for others.

      Shelby says "there's no linear path to software practitioner knowledge," and with so much information, constantly changing and deprecated, you can't expect people to have the same knowledge. Even senior people benefit from building from first principles, and Shelby tries to state up front who the intended audience is and what readers should know. Matty learns best when solving a problem, like struggling to understand Habitat until finding a use case, and likes "journey tutorials" that start with the problem and not just how to call a function.

      The "So What?" Question

      Shelby says developer advocacy makes you want to have experienced something before talking about it, and describes a second internship outbrief that the CEO wanted to attend, which led to coaching from people who present to three-star generals on organizational impact and the question "So what?" Shelby needed the "so what" on Kubernetes and load balancers when reading the SRE book in 2017, and says product pages and open source repos often fail to say what problem they solve or why to choose them over the standard library.

      Plan Your Message

      Matty says communication is a skill, and that both took a structured course called Communicate to Influence. Having a plan for what you're communicating beats a brain dump, which is what podcasts are for. Shelby says that when feeling insecure about being technical enough, the instinct is to drop a stack of books on someone's desk, but "the firehose doesn't really help people learn." Shelby is a stronger writer than ad hoc speaker, and writes internal docs in half an hour in response to questions. Both admit they rarely use the method they learned, though it's effective, and Shelby's manager suggested using it for CFPs.

      How to Learn

      Shelby's first tip is to figure out your learning styles, which are a spectrum: visual, auditory and kinesthetic. Shelby is a multi-modal learner who needs to interact with a thing every way and takes three to four times as long to grok something, but then can teach it. The second is the book Mindset, about growth mindset. Shelby grew up very hard on themselves and thinks the industry has gotten better about the gatekeeping and incredulity toward people who haven't heard of something yet, and Matty links a talk by Sasha Rosenbaum on Mindset from a meetup the week before.

      Matty ties it to psychological safety, which isn't just not being mean but includes how you answer a question in front of others. Matty recalls mocking a colleague in HipChat for a capitalized SSH flag, which sent the message that "If you make a mistake, Matt will make fun of you." You should think about the people listening, not the people you're talking to.

      Shelby says early in a career, teaching PhD-level aerospace engineers a Python library showed that everybody has gaps and "there's no linear path to expert." A platform engineer at Honeycomb admitted not fully grokking something, and Shelby blogs about things like breaking DNS on a Hugo site to share those gaps, and learns by thinking out loud and tweeting. Shelby also recommends Lichtenbergianism, a book about procrastination as a creative strategy, and embracing abortive attempts, the half-finished repos and unpublished posts. When learning a codebase, Shelby refactors and renames to see ripple effects, breaks things, then clears the index and starts a new branch for the real work.

      Risk and Privilege

      Matty notes learning in public carries different risks for different people, but cites Annie Hedgpeth, whose blog about learning InSpec became the documentation linked from the product page. Matty, who learned by doing without a degree, credits Scott Hanselman with the observation that someone who doesn't look like Matty wouldn't have gotten the same chance to not know what they're doing, and recalls being asked about a VPN in 1998 and told to put one in the next day.

      Shelby says luck mattered too: on the first day of a second internship, a product lead asked the linguist whether they wanted to write a translator, which turned out to be compiler theory. Shelby says being able to spin that into a positive story is not available to everyone, and that appearance and paper credentials lend credibility to some people. There's a disproportionate penalty for people who don't fit the mold when they raise organizational issues or push for psychological safety, and "Move fast and break things" doesn't work for everyone. Shelby thanks senior people and people outside minority groups for raising it, because it takes the burden off those more impacted and avoids being labeled as ruining everyone's fun.

      Matty plugs the return of Deserted Island DevOps on April 30th, with scholarships for underrepresented speakers who lack a Nintendo Switch or Animal Crossing. Shelby will speak at Developer Week in February and join an InfoQ panel.

      • Learning Stuff With Ali Spittel - ADO Episode
      • Managing Your Mental Stack - ADO Episode
      • Mindset: The New Psychology of Success
      • Sasha Rosenbaum's talk about Mindset
      • Lichtenbergianism: procrastination as a creative strategy
      • 58 min
      • The Six Plots of Tech Twitter

        Matty talks about tech Twitter with five guests, recorded in early January 2021: Sasha Rosenbaum, who had just joined Red Hat, Aaron Aldrich, also newly at Red Hat, Jeremy, head of DevRel at CircleCI, Quintessence Anx, a developer advocate at PagerDuty, and Kat Cosgrove, a developer advocate at JFrog. The idea came from Matty wondering whether tech Twitter has only six basic plots, like the six or seven basic story types credited to Kurt Vonnegut and Christopher Booker. Sasha, Jeremy, Quintessence and Aaron each posted lists, and Kat answered snarkily instead. The cold open is Kat: "Sometimes Twitter is not good."

        The Same Plots Again

        Kat nominates gatekeeping, in which a startup CEO or VC posts a terrible opinion about women in tech or bootcamp students, everyone quote-tweets to dunk on it for 48 hours, and the same thing happens again about four months later. Matty adds that anecdotes get treated as data: someone says it didn't happen to them and treats that as a refutation. Aaron says arguing past each other is a big trope. Sasha notes that telling a story about one customer is how sales and DevRel work, and Matty says the problem is when a personal anecdote is treated as refuting someone else's. Matty's example is that Kubernetes is bad because it fell over in one organization.

        The recurring topics include arguing about Friday deploys about once a month, whether developers should be on call, and front-end versus back-end, which Sasha says turns into gendered gatekeeping. Quintessence adds that Twitter works as a live status page of who's down. Aaron says that when the Friday deploy and on-call arguments die, the DevOps transformation will be done, and the periodic reminder not to deploy on Fridays inevitably tags Charity Majors and leads to the same argument with the same people.

        Clever, Funny and Hot Takes

        Matty says people who think they've come up with a clever joke should search for it, since the New Year's resolution joke has been made a thousand times. Sasha says not everyone checks Twitter 15,000 times a day, so every joke is new to someone. Matty, who studied improv, recalls the rule "don't try to be funny," and Kat says funny tweets are "intrusive thoughts." Matty says the carefully workshopped tweet never lands and the random thought gets the engagement. Sasha says Twitter downgrades a tweet that gets no likes in the first 30 seconds, and links to YouTube or Spotify, while native video does better.

        On defining shitposting, Quintessence says hot takes, and Kat says it's a non-serious statement taken to an extreme, never intended as true or factual. Jeremy and Aaron say it should be nearly void of valuable content, and Jeremy adds there's a purposeful-troll version where someone drops a statement and walks away. Matty notes a hot take that everybody agrees with isn't one, and Kat calls the often-argued ones, like washing cast iron with soap, "room temp" takes, since people still argue.

        When to Argue

        Sasha says arguing with people made Sasha angry and obsessed for a week, while not engaging means forgetting within hours. Quintessence only engages if someone said something awful to someone. Kat says if something is actively harmful, having an audience of almost 13,000 comes with a responsibility to say it isn't okay, but Kat won't engage with a crappy take about 10x engineers. Jeremy checks a replier's profile quickly to see whether they're in good faith. Matty frames it as public versus private: you may be arguing for the benefit of the people watching, and if nobody is listening it just makes you louder and angrier, so take it to a group text. Aaron offers three levels: call out bad takes publicly, have good-faith arguments loudly so others learn, and walk away from bad faith.

        Sasha says a blog post from another person with about 14,000 followers compared a follower count to a stadium of people listening, which stresses responsibility. Kat describes how a pithy tweet in a thread explaining a change to Kubernetes carried no weight at 4,000 followers, but 36 hours later at almost 12,000, people assumed Kat worked for Google or Docker, because Twitter has no context. Matty adds that as your audience grows, your tone and personal posts change, which isn't self-censorship. Sasha says putting "opinions are my own" in a bio doesn't help with screenshots, and Kat removed it since it's not a legal disclaimer. Matty says anyone who wants to go after someone won't read the bio, and shares a friend's email footer: "by opening or replying to this email, you agree that all of my opinions are correct." Sasha adds that Sasha's out-of-office replies are valid JSON.

        Tips for Using Twitter

        Kat's tip: if you're not a man and get a filtered DM from someone you don't follow that opens with just a greeting, don't respond, since often it's an inappropriate photo. Kat and Aaron add that nobody should open a DM with just a greeting, and to include the actual question. Sasha suggests that if you're not a man, consider not having your DMs open. Matty recommends private lists to curate about 100 accounts, since lists are always chronological and don't hit the algorithm. Jeremy suggests pinning lists in the mobile app. Quintessence says if your feed is monochrome, deliberately follow leaders who aren't white or aren't men, and Twitter's recommendations will follow. Sasha says to follow people beyond the big names, and that advice differs under or over 1,000 followers. Aaron says lists matter more when your follows act as endorsements. Matty and Sasha keep private lists of cute animal accounts for toxic days.

        Jeremy's last tip is to use your walk-away power and shut the app off. Sasha and Kat say to have at least one friend who understands Twitter and can take a rant, since many real-life friends only see screenshots on Facebook. Quintessence says a major news cycle mixed with sales pitches and Friday deploy takes can be stressful. Sasha reminds everyone that Twitter is also a community for support, hugs and cat videos, and on most days it's good.

        • The Six Main Arcs in Storytelling, as Identified by an A.I.
        • The Seven Basic Plots by Christopher Booker
        • 1 hr 1 min
        • Doing Releases Right with Scott Hain

          Matty talks with Scott Hain, a quality engineer at HashiCorp and a former release engineer, engineering services person, support person and engineer, about what goes into releasing software beyond a clever pipeline. The cold open is Scott imagining a customer's reaction to a flaky product: "They're like, oh, goddammit, fucking people. Why can't they make their shit work?"

          Delivery Versus Deployment

          Scott uses CD for continuous delivery, meaning making an artifact from code, whether a single binary, a binary with installers and wrappers, or something behind a SaaS. Continuous deployment is what happens when you actually ship a service or upgrade it. Matty says continuous delivery means your software is releasable at any point and shipping is a business decision, and cites Ken Mugrage's point that a holiday code freeze only freezes deployment, not the work. Scott adds that the release cadence is a human decision, and the aim is the best artifact sitting ready at any time: "if it's not ready to go, it's not ready to go, but it should be able to go." Matty notes "ready" doesn't mean complete, it means ready to make the decision to release.

          Scott says a release is also blog posts, announcements and other coordination that are mostly human-driven but can be automated.

          Versioning

          Once code is merged, the CI system creates a versioned, signed artifact, and the same code should yield the same binary, with caveats for Apple and Microsoft signing. Matty says it's a point in time, so you cut a new one, and versioning strategies stop you from ending up with "artifact.final.back.back.back." Scott says a rolling staged artifact removes the need for betas and release candidates, because customers could run it in their integration environments.

          On versioning schemes, Scott likes SemVer but says it doesn't always make sense, since customers ask about gaps, and another common option is a date timestamp. Matty jokes about putting a year in the major version, as with Office 2003, and notes the branded version differs from the artifact version. Matty says the wrong way is to be inconsistent, and complexity costs the people who must understand it. Scott disagrees slightly that the whole organization needs one scheme, as long as each project is consistent, because "consistency breeds automation." Scott also mentions meta-versioning for bundles of multiple artifacts.

          Workflows, Mandates and Tooling

          Matty says that if you try to find one true workflow for every project, nobody will love it, since compromise is when nobody's happy. Scott says mandates don't work and "do whatever works for you" doesn't either. A cross-functional release engineering or quality team can recommend tooling and set the things teams must follow, with exceptions allowed if explained, like SOX. Scott's key point is that engineering tooling has to make engineers happy, be easy to use and add value, or people will use workarounds.

          What Quality Means

          Scott says measures of quality differ by company: MTTR, incidents, number of bugs, Sev 1s, or uptime for a SaaS. Matty says people work to the metric you give them, telling Jez Humble's story of adding one test per sprint and getting assert equals true, and warns against tying metrics to compensation, which Scott strongly agrees with. Matty notes that if you measure Sev 1s, teams focus on "mean time to innocence," and wants metrics that are interesting and actionable, making your ears perk up. Scott suggests ISO/IEC 25010 (SQuaRE) and the Quamoco framework as starting points, both linked in the existing notes.

          Scott defines customer confidence as how much customers trust that your software does what they need, and how little they have to think about it. Matty recalls a tweet that customers want to be unaware your stuff exists and cites the Futurama line "sometimes when you do your job right, nobody even knows you did it at all." Scott says a telltale sign of low confidence is a customer saying "We're fucking tired of being your QA," and points to Nicole Forsgren's book Accelerate for why quality speeds you up.

          The Nuts and Bolts

          Scott says to make PRs run as many tests as is reasonable within a reasonable time, and to build an artifact from a PR so a developer can pull it down and debug locally. When merging, run exactly the same tests, because main branch creep is real. Build a workflow with tight feedback loops and quality gates in which each stage raises confidence: quick, high-value unit tests first, notifications, tests on the binaries, notarization, load tests, a long-running instance with customer-like data, and upgrade tests. Scott says to talk to customers, since a bad upgrade experience means people won't upgrade, and then you get a Sev 1 and a four-day upgrade. At the end, you have a releasable artifact in a staging area, and a button that makes it public.

          Getting There From Here

          Most organizations aren't greenfield. Scott says treat it as a migration in chunks with story mapping, noting that a year-long fix of a release process at one place was painful and that Scott isn't a fan of big bang. Matty agrees that any transformation should be iterative, since you'll miss some ifs anyway, and ivory-tower architects don't dictate everything. Scott recommends talking to engineers, without listening to everything they say, and to support staff, making a tactical plan for a small problem, writing RFCs, and tying it to company value and goals. Matty mentions a talk called Everything's a Product, applying product management to internal services such as release engineering, since your customers are the engineers, without the NPS score on a Jenkins pipeline.

          Scott's parting advice is to be empathetic to customers, engineers and teammates, and quotes Adam Jacob's line "happy people make happy software, which makes for happy customers," paraphrasing slightly.

          • Systems and software engineering — Systems and software Quality Requirements and Evaluation (SQuaRE) — System and software quality models
          • Operationalised Product Quality Models and Assessment: The Quamoco Approach
          • 49 min
          • 2020 Year-End Wrap-Up

            Joe Lahey edits the year-end wrap-up and opens it with a promise of 100 percent Kubernetes-free conversation, which lasts a few minutes. Matty, Trevor, Bridget, Jessica and Jeff pick favorite episodes, look back at 2020, and then spend the last third of the episode on Babylon 5. The cold open is Bridget: "I will try to apocalypse less in the future."

            Favorite Episodes and the Numbers

            Matty's two picks are Deserted Island DevOps, which Bridget had warned would have far too many people on one episode, and Breaking Down Gates with Tim Banks, which started as an ops conversation and "ended up talking about something else" that was better. Trevor and Jeff both pick the DevOpsDays Chicago 2020 episode, with Jeff calling the event "the best virtual event that I have attended." Jessica picks Don't Worry, Do Care with Aaron Blohowiak, about Netflix letting developers start as many services as they want: "don't worry about what it costs if this is worth it, but care how much it costs." Bridget picks Tea and Anarchy, for bringing together "overlapping, intersecting, yet disparate points of view."

            Matty's stats caveat is that there are "3 kinds of lies: lies, damn lies, and podcast listening statistics," so no numbers get shared. The most listened-to episode of 2020 by a wide margin was Deserted Island DevOps. Second was We're Always Learning with Patrick Debois, which was also the first episode with new co-host Jeff. Third was the communities episode with Jono Bacon. Matty's method for real listens is to cut downloads in half, and to treat downloads within 24 hours of publishing as a proxy for subscribers, since podcast apps download new episodes whether or not anyone listens. Both numbers keep growing.

            Pre-Recorded Versus Live

            Bridget has come to like the pre-record format, which let Joe edit the KubeCon EU Helm talk as a pop-up video with commentary from the other project maintainers. For KubeCon North America, the talk was a podcast-style conversation with no slides, and about an hour and a half of footage was cut to a 35-minute slot. Jeff adds that pre-recording lowers a barrier for underrepresented and nervous speakers.

            Matty describes having changed position on pre-recorded versus live after a long talk with Jessica, who argued that live lets speakers reference each other's talks, and then changing back. Matty ties it to work as imagined versus work as done: in a virtual event, people tend to "pop in, do their talk, peace out," and the one exception Matty saw was Deserted Island DevOps, where all the speakers sat in Zoom together all day. Jessica says it's the only virtual conference actually attended this year. The larger point from Matty is not to copy the physical event: for DevOpsDays Chicago the rule was "I don't wanna hear a damn word about technology," and to start from outcomes. Matty predicts a virtual buffet line will appear within two months, and Jeff calls the skeuomorphic approach a failure to take advantage of new avenues. Matty says to look at small events for innovation, since "the risk profile is less" and nobody has hundreds of thousands of sponsor dollars on the line, and Deserted Island DevOps was "Austin fucking around."

            What Happened in 2020
            • Trevor became a product manager, founded the Illinois Shuffleboard Association as its treasurer, learned green screening for the virtual DevOpsDays Chicago, and got a puppy on January 1st.
            • Jessica gave a keynote with Avdi at Codebeam on March 7th and 8th, closing the conference, and calls it the close of a conference speaking career, for now. Jessica now teaches workshops, including Invitation to Systems Thinking with Kent Beck.
            • At the end of 2019 Bridget announced a plan to travel less, and apologizes for "causing the apocalypse." Bridget also passed the global chair of DevOpsDays on to Matty after five years, saying it was time for the next generation of leaders.
            • Jeff finished the book Operations Anti-Patterns, DevOps Solutions, and describes the kids seeing their names in the dedication. Jeff's framing for 2020 is that "we're all in the same storm. We're not all in the same boat."
            • Matty's last in-person event was DevOpsDays New York, then a move from PagerDuty to Red Hat's transformation office, focused on state and local government, where the equivalent of "we're not Netflix" is "we're not the Department of Defense." Matty also started DevOps Party Games with Jeremy Meese, a monthly streamed game show built on custom Jackbox-style content, with a second league in a friendlier time zone planned for January.
            • Babylon 5

              Matty tweeted that listeners could ask the hosts anything for the year-end show and got one question, from Josh Zimmerman to Joe: what's the best episode of Babylon 5? The question traces back to an Ignite talk Joe gave at DevOpsDays Madison in 2016, on the most influential TV show no one had ever heard of. Joe's pitch is that it was one of the first shows with an overarching plot, running one story over five seasons. The short version, Joe says, is that it's Deep Space Nine "but good," which Trevor and Jeff push back on.

              Bridget tells of being asked in a job interview, "Star Wars or Star Trek? Show your work," and answering Babylon 5: Star Wars is fantasy, Star Trek is a utopian future, and Babylon 5 has a real future where dock workers strike and people are locked out of their offices for not paying rent. Joe's friend Christian Harrow replied on Twitter that the obvious answer is Severed Dreams, so Joe went with the best and worst episode of every season instead. The best picks:

              • Born to the Purple (season 1, episode 3)
              • The Long Twilight Struggle (season 2, episode 20)
              • And the Rock Cried Out, No Hiding Place (season 3, episode 20)
              • Moments of Transition (season 4, episode 14)
              • Day of the Dead (season 5, episode 8)
              • The worst picks are Believers, Confessions and Lamentations, Walkabout, Racing Mars, and Phoenix Rising, the season 5 episode about the telegoths. Bridget explains that the show was rushed to wrap up its plot when cancellation looked likely, then rescued by TNT with a fifth season that had nothing left to do, and disagrees with Joe on Racing Mars. The telegoths also get flagged for lighting candles on a station where earlier episodes treated oxygen consumption as a serious concern.

                The episode drifts from there into vehicle-themed 90s TV (Airwolf, Knight Rider, Street Hawk), the Star Trek books Trevor keeps buying, including one on the transition to a post-scarcity society, and the planned segment on looking forward to 2021, which Matty calls "an empty dock." The wrap-up ends with Trevor's projects (a Battlestar Galactica model kit), Among Us, Bridget's Hunt a Killer boxes, online trivia and three Dungeons and Dragons campaigns.

                Favorite Episodes
                Matty
                • "Breaking Down Gates" with Tim Banks
                • "Deserted Island DevOps" with a cast of thousands
                • Trevor
                  • "DevOpsDays Chicago 2020"
                  • Jessica
                    • "Don't Worry, Do Care" With Aaron Blohowiak
                    • Bridget
                      • "Tea and Anarchy" With Alice Goldfuss and Ian Coldwater
                      • Jeff
                        • "DevOpsDays Chicago 2020"
                        • What happened in 2020?
                          Trevor
                          • A year of PM land
                          • Streaming fun with Shuffleboard (and a little DevOps). Check out the Royal Palms Shuffleboard Shufflinsanity, or look at some shuff.io livestreams.
                          • Virtual conferencing!
                          • A doggo (Friday)
                          • Jessica
                            • Jessitron, LLC
                            • Cut my hair
                            • Last Keynote Codebeam March 7&8 last talk with Avdi
                            • systemsthinking.dev
                            • Bridget
                              • I planned to travel less. Sorry for causing (points to everything).
                              • Fewer events but playing with the pre-record format
                              • Jeff
                                • Got a new gaming table
                                • Wrote a book! Operations Anti-Patterns, DevOps Solutions
                                • Matty
                                  • Last in-person conference was DevOpsDays NYC
                                  • New job at Red Hat!
                                  • DevOps Party Games
                                  • Took over from Bridget as global chair for devopsdays
                                  • Did "less" speaking this year ("only" 9-10 talks)
                                  • 90s Sci-Fi Roundup

                                    Joe's Babylon 5 talk at DevOpsDays Madison 2016

                                    Joe's Top Babylon 5 Episodes
                                    • Season 1, ep 3: Born to the Purple, Great Londo episode.
                                    • Season 2, ep 20: The Long, Twilight Struggle. End of the Narn/Centauri War. One of the best FX shots of the whole show. One of the best Londo/G’Kar eps.
                                    • Season 3, ep 20: And the Rock Cried Out, No Hiding Place. More good Londo/G’Kar.
                                    • Season 4, ep 14: Moments of Transition. End of the Minbari Civil War. Good Delenn ep, better Neroon ep.
                                    • Season 5, ep 8 Day of the Dead. Duh.
                                    • Joe's Worst Babylon 5 Episodes
                                      • Season 1 ep 10: Believers. Basically Christian Scientists in space. I could have chosen TKO (Bloodsport in Spaaace!) or Infection ("He took a pretty bad hit!"). Believers is super-preachy and it has the pedigree of being written by David Gerrold (wrote the tribbles ep of original recipe Trek)
                                      • Season 2 ep 18: Confessions and Lamentations. A deadly plague threatens the Markab. Another weak Dr. Franklin ep. Also, a deadly plague, read the room B5!
                                      • Season 3 ep 18: Walkabout. Hey, look! It’s Wallace’s mom from Veronica Mars but oh, god, the songs!
                                      • Season 4 ep 10 Racing Mars. Marcus AND Dr. Franklin together on Mars! Nuts and gum, together at last. Honorable mention: The Deconstruction of Falling Stars.
                                      • Season 5 ep 11: Phoenix Rising. Fucking telegoths.
                                      • 1 hr 26 min
                                      • Breaking Down Gates with Tim Banks

                                        Matty talks with Tim Banks, recorded on Friday the 13th in November 2020, about gates and barriers in tech. The conversation opens with a chili argument from DevOps Twitter and turns into hiring, inclusion and how companies treat people during the pandemic. The cold open is Tim: "If you don't have black women that are rising through your ranks, you're fucking up."

                                        Chili as a Metaphor

                                        Tim has strong views on chili, which for Texas has two ingredients, meat and heat, with no beans and no tomatoes, and uses it as a metaphor. Chili, like DevOps, means different things in different places: is DevOps a job title, a methodology, a culture or a group? The same title comes with different responsibilities at different companies. Matty notes the metaphor breaks down because chili has a definition and DevOps doesn't, and adds that at a bank everyone has the title vice president, which means you're not in leadership.

                                        Standardizing Levels and Interviews

                                        Tim would like tech to have something like the journeyman levels in the trades, so people know where they stand, and Matty adds that electricians have unions and certification boards. Tim says that if we call ourselves engineers, there should be standards, whether a union or a certification. Standardization would help with treatment, pay and inclusion, and lower barriers by making hiring less arbitrary, such as trivia interviews where someone is asked for an inverted binary tree they will never write in the role. Tim also worries about startups expecting junior people to work 60 to 80 hours for no equity.

                                        Matty mentions Kat Cosgrove's All Day DevOps talk on gatekeeping, which showed a junior job posting that described an entire team. Matty says job descriptions get garbled between a hiring manager's nice-to-haves and an engineer who doesn't know how to write job requirements, and recalls keyword-stuffing a résumé for recruiter software, like listing every ProLiant model. Tim still gets Solaris contract offers because the résumé mentions it.

                                        Two Kinds of Gatekeeping

                                        Tim describes two gates: getting in, from résumé screen through interviews to the offer letter, and staying in, with promotions, raises, reviews and projects. Many companies measure inclusion with lagging indicators like hiring numbers, and Tim says to look at retention and promotion. Tim says it's not a pipeline problem: people leave, even if you get them in. A good gauge is an anonymous survey asking whether someone who was LGBTQ would feel comfortable coming out to co-workers and leadership, and the answer has to be an emphatic yes. Matty likens it to an NPS question because it has skin in the game.

                                        Matty compares counting hires to monitoring and sentiment to observability. Tim agrees: observability is insight into what's going on inside, where hiring and retention numbers are "when you get paged at three o'clock in the morning that something is down because it's already broken." Matty adds that you can't put a Nagios agent on your people, and that becoming a people manager is a career change, not a promotion.

                                        Being the First

                                        Matty worries about the difficulty of being the first person from an underrepresented group on a team. Tim says you can't hire a junior person to be the only one and need someone with more experience who can navigate it, and that you have to fix the culture and listen without defending. Tim says "when you have a more inclusive culture, you will have more inclusive hiring practices. Full stop." Matty says a failed first attempt is not a reason to stop, because you have to keep trying.

                                        Tim says the first hire is being asked to do their job plus fixing the culture and being a pioneer, which is three jobs, and so should be paid more, and the worst outcome is finding out you're a token. Tim recalls a company that announced a diversity push and then celebrated progress hiring more white women, and notes "the things that you measure are the behaviors you're going to incur." Matty describes a frozen middle: people who support diversity in principle and then stall when it gets hard.

                                        People in a Pandemic

                                        Tim says tech has often told people the job is typing on a keyboard, when people need to feel they matter as whole people. Tim says that in the pandemic tech is losing women, especially mothers dealing with kids at home, and it's "unforgivable." Matty adds that much of the loss is invisible. Tim says if a hiring manager loses a woman who has to take care of family, "you fucked up," not the person who left.

                                        Matty and Tim say companies have the money. Companies are saving on rent, commuting benefits and parking, so they could offer stipends or a one-off payment, like $1,000 each to 100 people, which is a rounding error for most VCs. Matty argues investors don't have a fiduciary responsibility to make shareholders insanely rich, and suggests that if growth is 300 percent instead of 1,000 percent, that's fine, and to invest in people. Matty ties this to resilience: organizations rebound from unmodeled disruptions like a pandemic through depth of capacity, which comes from people, "because it's banked." Tim says you must put into the granary in good times and take care of people all the time, since they'll leave "the second they can." Matty says the bar is so low that being slightly less bad than everyone else will make a company look like the best in the world.

                                        Tim ends by saying that without insight into what your people are going through, the only signal is a lagging indicator when something goes down. Matty adds that you'll see it a year or year and a half later, wondering where everyone went: "Here's your chance to get it right." Tim is @elchefe on Twitter.

                                        • Texas Chili Parlor in Austin
                                        • Chili John's - the chili place Matt's friend brought him to in California
                                        • Gatekeeping and the DevOps Revolution: We Haven't Always Known Everything - Kat Cosgrove at All Day DevOps 2020
                                        • Lending Privilege - Anjuan Simmons
                                        • levels.fyi/
                                        • Shout-out to Matt's friend Marcelo for the link for Chili John's (and for taking Matt there so many years ago)

                                          51 min
                                        • devopsdays Chicago 2020

                                          Matty and Trevor host a supersized panel about devopsdays Chicago 2020, the first virtual version of the event. The panel: Kevin Reedy, a technical account manager at Kong who ran the AV team, Sasha, a long-time Chicago organizer who pushed for going virtual, Jason Yee, Director of Advocacy at Gremlin, a speaker, Kat Cosgrove, a developer advocate at JFrog who moderated, Laura Santamaria, a developer advocate at LogDNA who gave an Ignite, and Chris Read, an organizer on the virtual team who has been involved since the first devopsdays in Ghent. The cold open is Jason: "With the Yak and with DevOps Deep Thoughts, everything else, it was very much DevOps Days Chicago."

                                          Deciding to Go Virtual

                                          The organizers took three to four weeks to decide, and Matty says the focus was on outcomes and not on platforms, since "that's so very DevOps," being outcome-driven and not implementation-driven. Matty says engineers default to "how would we do that in Slack?" and argues the wrong question about virtual events is how to replicate the hallway track, since that gets you "augmented reality VR expo booth nonsense." The right question is what you get from the hallway track. Sasha pushed for a virtual event so the community would have somewhere to gather, and for it to be free, with talks watchable on YouTube. They looked at many virtual platforms and said no to all of them. Chris says the worry was making sure participants and speakers felt they got value.

                                          The event was split between an AV team in a studio in Chicago, which recorded and streamed, and a virtual team handling participation, so Matty and Sasha, both in the studio, didn't see how the participant side went.

                                          Audio and Video

                                          Kevin says AV at devopsdays Chicago started out terrible and iterated upward year over year, including live captioning, but this year "we pretty much threw that entire thing out the window." Kevin's priority was quality, and Kevin insisted there be no live remote presenters over Zoom, since internet, audio and lighting are hard to control, and the first minutes of many virtual talks are AV problems. Presenters could record in a studio paid for by the event, with guidelines for two cameras, or self-record using a guide. Trevor wrote most of the self-recording guide, drawing on experience recording the podcast. AV Chicago, the local partner, did all the editing for consistency. After each talk, a live Q&A or fireside chat over Zoom, joined early so lighting and audio could be checked, ran on the main YouTube stream.

                                          Kevin says the studio was in a Chicago music venue Kevin loves, and the AV Chicago team was wonderful. Matty adds that the fireside chats were in front of a fake fireplace that Sasha figured out how to turn on, and that speakers could be in the chat during their own talks, which Matty says gave more riffing opportunity than in most virtual events.

                                          Why Discord

                                          Matty says the plan started with Slack, then Matty wished Slack had video channels like Discord's, and realized Discord was the answer. The event was low-risk since it was free, so they could take a gamble. The main concern was abuse, and Discord offered more granular permissions and moderation than other tools. They used bots so people had to accept the code of conduct to get access. Matty recalls Dr. Richard Cook asking what they'd do if everything went wrong on Discord, and the answer was turning it off, plus recruiting volunteer moderators from across the devopsdays community worldwide. Matty also says they recorded how-to videos and over-rotated on onboarding. Matty is writing a blog post with details.

                                          Kat has done roughly 25 to 30 virtual conferences this year and calls this one of the favorites, since Discord allows friendly, genuine communication with attendees and speakers, is easy to moderate, and "if DEF CON can pull it off with tens of thousands of attendees, than anybody should be able to." Kat says moderating was often facilitating, and once a good question started, a room self-managed for 15 or 20 minutes, feeling as close to an in-person conference with real breakout rooms as Kat has had.

                                          Breakouts Behind the Scenes

                                          Matty didn't call them open spaces, but they followed the spirit, with attendee-suggested topics in text or video rooms. Chris says the proposal channel went unnoticed at first, so moderators reseeded topics, people voted with emojis, and video rooms were capped at 25 users, so a moderator had to be in place first. In the first round, people hopped in and out and nobody turned on cameras, and it took five or six minutes to get people talking, but it got easier. Organizers and moderators had separate channels, Jerry acted as ringmaster for the virtual team, and a read-only FAQ was updated live. Chris didn't hear about any problems getting into Discord, joking "We didn't observe it, therefore it did not occur." Chris found the virtual side more draining than an in-person conference, with constant context switching.

                                          The Speaker Side

                                          Laura recorded the Ignite at home, following Trevor's guide, with a bedsheet taped to a coat rack to reflect light, and needed around 15 takes because of cars passing the window and a hand-off of an inflatable microphone. The Ignite speakers passed inflatable mics from one to the next, and Laura dropped the mic onto a pile of blankets that the dogs then lay on. Matty ordered the mics in sets of five with random colors, and had hoped someone would mess up the hand-off. Laura says watching your own talk while chatting is more nerve-wracking than live, but being in the discussion all day meant nobody missed the hallway track.

                                          Jason recorded in a studio near train tracks in a sketchy area, ended up sitting, and was the last talk of the day, followed by a fireside chat with Matty and the DevOps Yak. Jason says the chat let Jason skip watching the video of the talk, and that Matty and Sasha were exhausted, which Jason says shows that virtual events are a lot of work. Sasha says Zoom-style events are more exhausting because you don't get five minutes chatting with a friend and a coffee. Jason says a polished recording helps with future CFPs, and jokes that the bar is now high for organizing.

                                          The Yak, Deep Thoughts and Stickers

                                          Josh Zimmerman traditionally gives a funny Ignite, and with extra stream space Josh recorded a series of DevOps Deep Thoughts that played through the day. Kevin says sponsors got a three-minute pitch, longer than usual, sandwiched between two Deep Thoughts so people would not skip them. The team also filmed the yak costume against a green screen and Trevor turned about 25 minutes of footage into yak bumpers, which Matty says led many viewers to think the yak was live. Sasha says they were a major improvement over running dry content.

                                          Matty had avatar stickers of speakers and organizers made, designed by Kelly Mahoney, in place of the professional headshots of the previous year. Laura says it was the best speaker gift ever, because it was a drawing of Laura. Jason says it worked because it was something nobody thought to buy. Kat, who had never been to devopsdays Chicago, says the small things going wrong, like shuffling moderators and scrambling for topics, made it feel more like a normal conference than platforms that hide every hiccup, and felt "more normal for a day, which is extremely valuable right now."

                                          Matty ends by reading part of an iTunes review from a nurse who finds the podcast gold, and says reviews like that are why the show has kept going for seven years.

                                          • Matt's blog post about "howto"
                                          • Rich Burroughs’s wrapup post
                                          • https://matty.wtf/yak-wtf
                                          • https://twitter.com/SoSplush
                                          • Love, Yaktually video
                                          • Recording of the event livestream
                                          • 1 hr 8 min
                                          • Tea and Anarchy with Alice Goldfuss and Ian Coldwater

                                            Bridget talks with Ian Coldwater, who lives in Minneapolis and specializes in hacking and hardening containers, Kubernetes and cloud-native infrastructure, and Alice Goldfuss, who lives in Portland, Oregon, and has a background in site reliability engineering, systems programming, software engineering and network engineering. The conversation covers container security, CVEs and disclosure, privacy boundaries for people with public profiles, and tea. The cold open is Alice on disclosure manners: "I think it's supposed to be bad manners to just be like, lol, your shit's broken."

                                            Containers Are as Secure as the Stack

                                            Ian says container security has to be thought of holistically, because containers share resources with each other and the host, so "your containers are as secure as your stack is," including silicon, operating system, kernel and what runs in them, and defense in depth matters. Alice has never worked on a container team with a dedicated security person, and says security is usually an afterthought, a checkbox before shipping. People often choose containers for security, such as running customers' arbitrary code, while security people say an insecure box means insecure containers, possibly more so with more ports open. Alice warns about false prophets and marketing, and says if you pick containers for security, you need an expert like Ian and must implement what they say.

                                            CVEs and Disclosure

                                            Ian explains a CVE as a taxonomy: someone who finds a vulnerability submits it to the CVE Numbering Authority, it gets a number, and you can look the number up in a large database. The numbers aren't memorable, so people name their vulnerabilities. Ian recently got a first credited CVE with a group, a credential leak in containerd 1.2.x, and mentions an earlier Kubernetes CVE for which friends published a proof of concept that "honked Kubernetes to death."

                                            Alice says when a CVE lands, you need an inventory of the versions you run, since "otherwise, you find out about them on Twitter." Then ask whether you run the affected version, how likely and severe it is on your fleet, and how much of your infrastructure is affected, and set a mitigation deadline. CVEs typically aren't announced until a patch is in the works, and an older version might make patching gnarly, which is one reason to keep doing rolling upgrades.

                                            Ian says responsible disclosure is a matter of some debate. At best you write to the security contact, get a friendly reply, file the CVE through maintainers, who set severity, and agree on a timeline. Sometimes vendors threaten to sue the finder, which is bad behavior and how "you get 0-days dropped on Twitter on you." Ian notes that someone who reports a bug is showing good faith, since they could sell or post it. Alice adds that if you get an email saying someone found a vulnerability on your site, do not respond with threats, because that person is trying to help unless the email continues with demands for payment. Bridget's analogy: "hi neighbor, your window's unlocked."

                                            Below the Software

                                            Alice says the kernel is software too, and has a well-entrenched maintainership that gets patches out, while operations teams are used to patching it. Hardware has its own issues. Alice recalls an F5 announcement a couple of hours before a leap second that some load balancer versions had a vulnerability triggered by it, which meant interrupting an important meeting with executives. Alice also describes the "hotel maid" scenario of a device being placed in a laptop and security researchers scanning machines before and after travel.

                                            Public Life and Personal Security

                                            Bridget asks where they draw boundaries between public work and private life. Alice takes a physical safety mindset, doesn't tag locations, shares restaurant visits only hours after leaving, and has been recognized from Twitter on the street and in restaurants. Alice is a protected voter in Oregon and currently doesn't share an employer on Twitter, because the boundary has been good for mental health. Alice enforces parasocial boundaries, inviting approach at events but not at a restaurant or crossing the street. Bridget posts about work because it fits open source, and says familiarity doesn't mean friendship.

                                            Ian says everyone has their own threat model and comfort level. Ian tells of a single week when a parent at a kid's karate class and the person at a bodega counter both announced they followed Ian on Twitter, and Ian decided to accept being visible. Ian's advice: "err on the side of not being creepy," and "if you know for a fact that you're being creepy, maybe stop." Alice adds that constantly being watched has paid off in professional settings as a kind of trial by fire.

                                            Tea and Geese

                                            Asked for the most delicious tea, Alice says it depends: coffee drinkers may like smoked teas like Lapsang Souchong or malty ones like Assam, people who like fruity flavors might try an oolong, and people who like savory might try a Japanese green like a Fukamushi sencha. Alice's recent favorites include a Weishan Bao Zhong and a Dan Cong oolong that smells like currants. Ian isn't a tea snob but enjoys friends' tea.

                                            The goose thing came from the video game Untitled Goose Game, which Ian saw as an allegory for hacking because the goose chains together innocuous objects to exploit them. Independently, organizers of the Kubernetes Contributor Summit at KubeCon made a goose-themed CTF where GitOps makes a stuffed goose honk, and Ian's keynote was about it. Now "lots of Kubernetes people honk at each other."

                                            Making the Year Better

                                            Bridget is trying to elevate voices other than Bridget's own. Alice says impact depends on bandwidth, and changed the PDX DevOps meetup, a group of about 60, to go online and host topics beyond technology, including ethical organization at a company after the George Floyd protests, resources for protesting safely, stress release and emergency preparedness after wildfires. Ian says that as someone with access to tech community resources, the answer is redistributing money and resources to BIPOC youth doing organizing on the ground. Alice agrees.

                                            Image credit: Tea and Anarchy, modified from Anarchist Revolt

                                            Font: 1403 Vintage Mono Pro by Jeff Kellem

                                            38 min
                                          • State of Open Source Security with Alyssa Miller

                                            Matty and Jessica Kerr talk with Alyssa Miller, an application security advocate at Snyk, about findings from Snyk's annual State of Open Source Security report. Alyssa has been in security for 15 years, started as a hacker at 12, and is a self-described recovering developer who spent most of a decade in financial services. The report combines open source data from GitHub, GitLab and Bitbucket, aggregated data from Snyk's own product, and a yearly survey of practitioners, developers and ops people. The cold open is Alyssa's line: "No, no, no. Gates break DevOps, period. You can't do it."

                                            Dependencies Behind Dependencies

                                            Alyssa says the number of packages keeps growing, with npm's count nearly doubling every year, and that most vulnerabilities come from indirect dependencies, not the ones you chose, especially in Java and JavaScript. An example from the report is an 80-line JavaScript app with 7 dependencies that expands to 59 more and turns into 750,000 lines. Matty asks whether to fix this or accept it, and Alyssa says we need awareness and tools because it isn't going away: security has preached "don't roll your own encryption," and reuse was the panacea when Alyssa was a developer. Alyssa raises the software bill of materials, noting an FDA advisory about a popular open source package that medical device makers couldn't answer for, because they didn't know what was in their software, much like Heartbleed.

                                            Jessica says that if you write it yourself you'll still have vulnerabilities, only nobody sends you mail about them. Matty cautions that open source might have been reviewed, not that it was, and Jessica says nobody looks at what Jessica publishes on npm. Alyssa says package health has no generally accepted measure, but popularity, age, active maintenance and accepted PRs give an idea, and popular packages get more scrutiny now and later. Academic researchers, such as a security lab at UC Santa Barbara, report vulnerabilities to Snyk after scanning thousands of projects for patterns.

                                            What Improved

                                            Alyssa says the total number of new vulnerabilities grew more slowly than the year before, with fewer reported in 2019 than in 2018, which Alyssa calls a positive indication but "not ready to say, hey, we're getting better at security." On who's responsible for application security, about 85% said developers both years, but security rose from 23% to 50-55%, and operations went from barely registering to about the same, which Alyssa reads as awareness that DevSecOps needs all three. More organizations review their YAML and JSON and audit production clusters, though 31% said they didn't know or weren't doing anything, and 44% of respondents use Kubernetes.

                                            The Scatter Plot Surprise

                                            New this year was a scatter plot of vulnerabilities reported versus projects impacted. Cross-site scripting had many reports but few projects affected, while prototype pollution and deserialization had few reports but wide impact, including a Lodash vulnerability in 2019. Nothing landed in the upper right quadrant, which Alyssa reads as a sign that big, popular projects have eliminated the common flaws, while newer attack vectors have the big impact.

                                            Official Images Aren't Safe by Default

                                            Alyssa calls this the "stranger danger" story and says it was personal. A blog claimed official Docker Hub images have been scrutinized, but the report found the top 10 official images still had many vulnerabilities, with the Node image off the charts. Alyssa pulled the full Node image with 642 vulnerabilities, versus about 53 with the slim image, and compares it to the old server advice to minimize the operating system. Jessica says the full image is handy for development but production is a different animal, and Matty points out the temptation to go back to the image that works. Alyssa notes Docker has been adding scanning and higher-scrutiny image programs.

                                            SnykCon and Threat Modeling

                                            SnykCon is a virtual conference on October 21-22 with vendor-agnostic talks, keynotes such as Wendy Nather, and a fundraiser for the Bill and Melinda Gates Foundation. Alyssa will do a workshop on threat modeling in DevSecOps.

                                            Alyssa says threat modeling answers "what could possibly go wrong?" Traditionally it's a heavy process of data flow diagrams taking days, fine for waterfall but not sprints. Alyssa's approach adds a continuous-improvement CI: threat model each user story, where someone from the business can guess what attackers would want to steal, expose or deny, and that informs coding, test cases, automated tools and production monitoring. You won't make the system "unhackable," but you get incrementally better. Matty adds that monitoring is testing with a time dimension, and Jessica says talking to domain experts about what it shouldn't do gives a better understanding of what it should. Alyssa points to a Puppet survey finding that collaborative work like threat modeling builds more confidence in security posture than siloed pen tests and scanners, and says it helps prioritization: which vulnerabilities protect the crown jewels, and which are exploitable at all.

                                            Gates Break DevOps

                                            Alyssa has attended no fewer than 30 talks on DevSecOps that put quality gates between stages. Security has to be integrated in each phase, since "If you push motion in the pipeline back to the left with the feedback from a gate, you just broke DevSecOps." Matty adds that people who can't get through a gate figure out how to go around it, and tells of expensive security tools with few licenses that act as gates themselves. Alyssa cites a study where 85% of organizations said they'd pushed known vulnerabilities to production and 54% of those said it was to meet a timeline. So you have to accept that software ships with vulnerabilities and shorten the feedback loop. Jessica adds that gating slows security releases, since "change is on our side."

                                            Alyssa's favorite t-shirt says "Unhackable?" with "Here, hold my beer" beneath, linked from the show notes.

                                            • Snyk's State of Open Source Security report
                                            • SnykCon is coming on Oct 21-22! Register now!
                                            • Guess what you can threat model in devsecops! More about threat modeling in Pushing Left With Tanya Janca
                                            • Alyss's awesome t-shirts
                                            • 56 min
                                            • Incident Retrospectives with Amy Tobey, Alex Hidalgo, and Rein Heinrichs

                                              Matty talks about incident retrospectives with three people who care about learning from incidents. Alex Hidalgo, an SRE for about ten years, has a book coming out from O'Reilly, Implementing Service Level Objectives. Amy Tobey is a DevRel and Staff SRE at Blameless who started in tech around 1999. Rein Heinrich is a principal software engineer who helped make Puppet in 2009 and co-hosts the podcast Greater Than Code. The episode started on Twitter, where Alex and Amy were discussing retrospectives and Alex suggested they go on the show. The cold open is Alex: "Oh shit moments are just about my favorite."

                                              What a Retrospective Is For

                                              Matty defines the topic as what happens after service is restored, however you name it: postmortem, after-action review, retro. Amy frames an incident as an unplanned investment, with people time, software and cloud spend going in, so the question is whether the organization got the most out of it. Rein adds the view that an incident is an encoded message the system is trying to deliver, and while responding you decode just enough to restore service, so skipping the rest wastes the investment. Alex calls retrospectives the most sophisticated end of responding to a ticket, where you fix a problem in the best possible way and not just click close.

                                              Matty says during an incident the goal isn't fixing the problem but restoring service, and that the two halves depend on each other: you can only skip the rabbit holes during response if the organization has a social contract to decode the message afterward. Amy describes different paces of engineering, with incident response at the fastest and a retrospective on a much longer timescale, like the architecture phase. Rein compares this to Kahneman's thinking fast and slow, where fast is about performance, not long-term learning.

                                              Start Before the Meeting

                                              Matty notes the irony of "thinking slow" in a one-hour meeting, which should be a jumping-off point. Alex says you can start the slow learning during an incident with the incident command system, used conceptually, by asking someone to start the incident state document or the retrospective while responders focus on mitigation. Amy adds that assigning the scribe role is also a way to keep a nosy manager busy. Matty cautions with Ron Swanson that you should not half-ass two jobs, and Alex agrees that it takes a practiced organization where everyone knows who is doing what. Amy says if you show up to the meeting and most of the analysis isn't done, the meeting is a waste.

                                              Rein says the idea that learning happens in one hour is a little silly, since learning is happening in a dozen or more brains for days and weeks, and the meeting is for the things only possible with those brains in one room.

                                              Action Items and Their Deadlines

                                              Matty disagrees with a line from the PagerDuty postmortem guide that the most important outcome of the meeting is consensus on action items, though not if action items include questions to investigate and not just Jira tickets. Amy and Alex say that in the real world, follow-ups are the main point for most SREs, because they are the easiest to tie to business value and executives and directors want them. Rein's example is a junior SRE paged five times a week for the same thing, who won't accept "we're going to stop worrying about action items."

                                              Matty warns against SLAs on action items, such as completing everything within two sprints, since people will only agree to items they know they can finish, and an engineer should be able to come back and say the plan changed after looking closer. Amy's workaround is to get follow-ups into a prioritization process and then let go, turning choices over to the engineering and product teams. Matty says the people who prioritize work should be in the retrospective. Alex says many organizations lack buy-in, and Matty replies that nobody has it everywhere and change has to happen in both directions. Amy and Matty also note that when people manage the numbers they will game them, the Pareto-inefficient Nash equilibrium problem: people work to the numbers you give them.

                                              Narrative and Timelines

                                              Alex's goal is always to tell a story, since "we're storytellers," and finds a timestamp table less useful than a narrative of what happened first and next. Amy disagrees on timelines: the timeline is the outline before writing the narrative and common ground with readers, especially for complicated incidents. Matty calls the timeline supporting information, and notes that not every line in the Slack channel is worth including. Amy supports cranking out shallow incident reports cheaply and ubiquitously. Rein says there is no single timeline, with 12 people in a channel there are 12, and asking people to compare theirs is where the richness comes from. Matty adds that stories are more memorable than log entries, and that people don't read retros from other teams, which are the ones they most should.

                                              Matty says one of the biggest anti-patterns is only doing postmortems for Sev 1s, and Alex has been on teams where every page got a retrospective, even if it meant deleting the alert. Amy notes organizations that aren't ready to hear the reports, where small insurrections matter more, and Matty argues that writing short reports helps engineers learn to speak in business value.

                                              Making It Easier to Try

                                              Rein's rule is to ask what would have to happen to make a change easy, and suggests a Goldilocks zone between big, scary incidents and small, boring ones, and getting an organization used to trying things first, which can take six months. Alex has had success with facilitator rotations where people sit in on teams on the opposite side of the company, and Matty adds that a good facilitator isn't invested in the content, and that management facilitating is a problem. Matty says to stack the deck for change by starting with people who are interested, since they'll sand the rough edges.

                                              What They Changed Their Minds About

                                              Alex used to believe timelines were crucial, and now thinks the narrative is the important part, though they remain good starting points. Rein says timelines are what let you reinstantiate context in cognitive interviewing ("it was Friday, it's 9 PM"). Rein also no longer believes the hour in the meeting is the most important part, which Amy also said, and Amy no longer favors an independent meeting per incident, preferring the weekly incident review run at GitHub, where everyone came for a cadence of caring about incidents and people shared what happened in narrative form, with follow-up done out of band.

                                              Better Questions

                                              Alex suggests templates with prompts such as "where do we get lucky?", who happened to be online, and how to make sure we don't have to be lucky. Matty adds a prompt asking what questions aren't on the template. Rein suggests asking how priorities should change as a result of the incident, since action items tell people what to do without why they should care. To choose which incidents to study deeply, Rein suggests looking for cues like confusion, surprise or frustration, and in the mundane looking for the surprising. Matty's closing advice: find the mundane in the interesting and the interesting in the mundane, write good narratives, and "ask why 10 times, because if 5 are good, 10 must be twice as good."

                                              Alex's book - Implementing Service Level Objectives: A Practical Guide to SLIs, SLOs, and Error Budgets

                                              59 min

                                            About Arrested DevOps

                                            From the publisher's feed

                                            Arrested DevOps is the podcast that helps you achieve understanding, develop good practices, and operate your team and organization for maximum DevOps awesomeness.