Arrested DevOps

Arrested DevOps

By Matt Stratton, Trevor Hess, Jessica Kerr, and Bridget KromhoutTechnologyTech News
Download on the App Store

Arrested DevOps episodes

  • Don't Worry, Do Care with Aaron Blohowiak
    Cost Compression

    Aaron discusses the three big projects he's working on. The first is cost efficiency, or, as Netflix thinks of it, "cost compression."

    Jessica: "Oh, so you have to achieve cost compression without telling engineers not to spend money."

    Aaron: "If you don't believe you can predict very well, then what you should do instead is get really good at reacting."

    The panel discusses the difference between autonomy and agency.

    The panel talks about the utility of data dashboards.

    Aaron: "The dashboards are more like cost debugging ultilities."

    Access Isolation

    Aaron talks about his second big project, "a unified strategy for access isolation."

    Jessica: "You want the cells to have access to the other cells in the same muscle tissue, but if they need a nerve ending they have to say so!"

    Regional Growth and Availability

    Aaron explains the third big project he's working on, "Netflix's regional growth and high availability story."

    Aaron talks about how Netflix produces original content all over the world, and the coordination required to serve all the computation needs of the various projects in different places.

    Shared Reading List

    Ecology, the Ascendent Perspective

    Thinking in Systems

    Aaron's blog post, The Sufficiently Smart Engineer

    55 min
  • Service Mesh with Michelle Noorali and Delyan Raychev

    Bridget talks about service mesh with Michelle Noorali, a software engineer on Microsoft's Azure containers upstream team and a core maintainer on projects including Helm, and Delyan Raychev, a software engineer on the Azure networking side who has spent about a year and a half on reverse proxies and Kubernetes. The occasion is the launch of Open Service Mesh, announced the day the episode was published. Both guests say the smallest change worth a pull request is a typo, a doc fix or a clarifying comment, and Delyan loves leaving to-dos. The cold open is Delyan: "SMI seems to be like the lingua franca of service meshes."

    What a Service Mesh Is

    Michelle gives the textbook definition: a dedicated layer of infrastructure that helps you manage, secure and observe service-to-service communication. The problem is that in highly dynamic environments, where pod IPs change as things come and go, networking needs to be more dynamic as well: traffic encryption, access control, which service can talk to which, traffic shifting from one version of an application to another, and observability metrics.

    Delyan frames it from the point of view of a CTO who wants observability, security and traffic management but has busy engineers: run an install command and "you get all those extra features." In the past this came from libraries tightly coupled to a language, such as Twitter's Finagle, Netflix's Hystrix and Google's Stubby. A service mesh instead bundles a sidecar reverse proxy with your workload and pipes traffic through it.

    Is It a Man in the Middle?

    Bridget asks whether intercepting traffic is a man-in-the-middle attack as a service. Michelle says it's meant to prevent them. A common requirement, especially in enterprise and government settings, is mTLS, mutual TLS, where both client and server prove who they are, and it's nice not to handle that in code. Delyan lists three components: the reverse proxy sidecar, a certificate used to encrypt and decrypt traffic, and the control plane that tells the sidecar what to do. All three are open source, so you can review the code, and traffic leaving the sidecar is encrypted and flows only to services explicitly permitted.

    Do You Need One?

    Michelle says "You don't need a service mesh unless you need those things," and it suits environments with lots of microservices and specialized teams, such as Twitter or Lyft, which Kubernetes and containers made possible. Delyan adds that operators get lost in Kubernetes complexity, and a mesh controls east-west traffic between namespaces, helps with zero-trust networking among teams that don't trust each other, and gives auditability of which services exist and who talks to whom. Michelle explains that east-west is service-to-service traffic, and north-south is external traffic coming into the cluster. The data plane is the set of proxies carrying user traffic, and the control plane is the source of truth that configures proxies and manages certificates. Michelle notes the sidecar approach isn't the only one, since some meshes run a proxy per node. Delyan says the data plane must run nonstop with minimal latency because customer data flows through it, while the control plane has more flexibility for upgrades.

    SMI and Open Service Mesh

    Michelle says SMI, the Service Mesh Interface, is a set of APIs representing the most common functionality people want from a mesh, so that tools can build against a consistent API regardless of provider. Delyan compares it to a shared language invented by college friends from different Eastern European countries: you can try mesh A, then mesh B, without changing your policies, because SMI stays in the cluster and only the data plane swaps. SMI describes the topology, the control plane ingests it and tells the proxies what to do.

    Open Service Mesh is a lightweight, Envoy-based, Kubernetes-native, SMI-compliant mesh. Michelle says SMI covers mTLS, traffic shifting, access control and metrics, and OSM chose Envoy for community momentum and WebAssembly extensions. Delyan says the goals are source that's simple to understand and contribute to, effortless to install, painless to troubleshoot and easy to configure with SMI.

    The design philosophy is "no cliffs." Delyan uses the analogy of service meshes as motorcycles in a garage, each fine-tuned for different uses, so there's always room for one more. SMI doesn't cover everything in the proxy, such as circuit breaking and back pressure, so when you hit that cliff there's a dirt road instead: switch to XDS, Envoy's own configuration protocol, which is harder but lets you fine-tune the proxies.

    Getting Involved

    Michelle points to the GitHub repo's install guide and demo, which work on a local cluster such as Kind or minikube, and says feedback via GitHub issues or Slack is welcome. Bridget adds that a newcomer reporting that the walkthrough did not work issue is valuable because the authors can't see what they've assumed. Delyan hopes people enjoy reading the code and rename variables to make it their own. Michelle is @michellenoorali on Twitter, and Delyan is @DelyanRaychev.

    • SMI
    • Open Service Mesh
    • OSM logo art credit: @flynnduism

      36 min
    • Developer Experience with Stephanie Stimac

      Matty, back after a gap between episodes, talks with Stephanie Stimac, a program manager on the Microsoft Edge Developer Experiences team, about what developer experience is and why it matters. Stephanie has a web design degree, spent four years at an agency as a designer and front-end developer, and got a DM on Twitter from a PM on the Edge team looking for a designer in a PM role. The first three years at Microsoft were a hybrid of designer, front-end developer and PM, including design on the open-source tool webhint and a brief stint refreshing Chromium DevTools to look like Microsoft DevTools. The cold open is Matty's line: "I can make stuff up. Trust me. I'm good at that."

      What Developer Experience Means

      Stephanie sees developer experience as a subset of user experience design: the experience developers have when using your product, which for Stephanie is a web browser. Edge Developer Experiences brings together the DevTools team, the web apps and PWA team, and the Ecosystem team, which Stephanie is on, after the move to Chromium. Stephanie explains that the Ecosystem team works like developer relations, helping other teams scale up and get good documentation out, and works closely with the HTML platform team on standards and features such as CSS requests. Stephanie describes WebView2 as a way to embed HTML, CSS and JavaScript in native applications, and has seen it demoed in Excel.

      Matty compares it to the days of throwing things in Notepad and putting a green border on every div, and says tools like Firebug were revolutionary.

      Ask Developers What They Want

      Stephanie says the biggest thing is that the Edge team now comes "from a place of humbleness" and asks developers what they want. Stephanie understands that building a browser used to be more closed off, with an assumption that browser makers are web developers and so know what developers want, which isn't true. Focusing on your developers is "the key to building a great developer experience," because "if you can build some cool feature, but if no one uses it," it doesn't matter. Matty adds the danger that when you think you're close to your users, you're really orthogonal to them, and skip the research.

      Documentation, Support and Crisis Design

      For products where developer support is a sidebar, Stephanie says two things to bake in are documentation and support, which "can make or break a developer's experience." Documentation tends to be left to the end, with an assumption about what users know. Stephanie calls out some popular static site generators where debugging leads to an endless loop of docs with nobody to contact. Stephanie recalls an Eric Meyer talk about designing for users in crisis, the idea being that an experience a user in crisis can navigate will work for everyone, which applies to developers whose site has broken. Stephanie's advice is to keep iterating on documentation and stay open to feedback.

      Matty ties it to empathy: things make sense to you because it was your idea, and swagger output isn't documentation, since examples matter and "I am coming to solve a problem." Matty adds that if someone keeps asking how to do the same thing, "that's on you," and compares it to learning to drive a manual so you can drive an automatic.

      The Web We Want

      Stephanie spends about 60 percent of the time on The Web We Want, a cross-browser and standards initiative that started on the Edge team but isn't Edge-specific. It's a forum for developers to say what's missing from the web platform, asking what they'd change if they could wave a magic wand. So far it's had about 150 valid feature requests or gaps, and "developers are really hungry to give their feedback." HTML controls is one request that matched work already underway. Some submissions are things standards groups decided years ago weren't worth the investment, and now there's data from developers.

      Matty links it to the saying that in open source "no is temporary, yes is forever," and that for no to be temporary you have to keep looking. Matty adds a change-management tip: reassure people that with the information they had, they made the right decision, and now things are different.

      Design Skills and Tracking the Work

      Stephanie says "at my core, I am a designer," solving problems and looking at the whole developer experience, such as what a developer sees when arriving at the website. Stephanie gave design feedback on the Grid tooling going into Chromium, since a designer debugs layout differently from someone who only develops.

      Stephanie says the team tracks engineering work in Azure DevOps, and that Web We Want submissions and problems extracted from interviews are hard to track there, because it isn't an engineering task you can give a number of dev days, and there's a "bucket of wants." Stephanie doesn't have a solution, noting the tool was built for dev work. Matty says you'd have the same problem in Jira, and it's the classic DevRel problem of tying work to value. Stephanie adds that features in DevTools aren't viewed as done when they ship, since usage and feedback continue to drive iteration, and Matty says organizations' measures don't map to continuous improvement.

      Empathy and History

      Stephanie's steady message is empathy: talk to a subset of your users about their pain points and what they like, and "embracing your empathy and shed your assumptions." Stephanie likes telling stories about history, and in the HTML Controls talk dug into a 1994 or 1995 specification. Stephanie says an unresolved developer complaint can linger for years, and calls Internet Explorer a great example. Stephanie is speaking about HTML controls at FrontCon in Latvia the next month, and has a YouTube channel with the February version.

      • The Web We Want
      • Webhint tool
      • Designing For Crisis - Eric Meyer talk
      • FrontCon - upcoming speaking appearance for Stephanie
      • Stephanie's current and past talks
      • Stephanie's talk on web controls
      • Go to Stephanie's YouTube channel for past talks!
      • 50 min
      • Deserted Island DevOps

        Matty talks with the organizers and speakers of Deserted Island DevOps, the DevOps conference held inside Animal Crossing a few weeks earlier, in what Matty calls the most simultaneous guests the show has had in an episode it will release. The organizers are Austin Parker, a developer advocate at LightStep, and Katy Farmer, a free-range developer advocate. The speakers are Ian Coldwater, a lead platform security engineer at Heroku, Dave Sudia, a senior DevOps engineer at GoSpotCheck, Kat Cosgrove, a developer advocate at JFrog, Jacquie Grindrod, a developer advocate at HashiCorp, Mia Moore, a developer advocate at IBM, Adrienne Tacke, a developer advocate at MongoDB, and Aaron Aldrich, a developer advocate at LaunchDarkly. Several lines in the transcript are labeled only "Participant," so this summary attributes only what's labeled. The cold open is Austin: "It's amazing what you can do in 30 days and someone with 50,000 Twitter followers."

        How It Happened

        Austin says it happened because Austin and Katy were bored. With COVID, they started Twitch streams, and Katy's idea was a Friday stream about something other than technology, which began with Animal Crossing. Austin joked on Twitter about building a trade show booth in Animal Crossing, someone suggested a conference, and with April 1st the next day Austin put up a web page, planning to call it an April Fool's joke if nobody cared. There were a hundred signups in a day for an event with nothing behind it, and Katy remembers a message saying "I made a joke, but now it's real." Austin calls that the hallmark of a good idea.

        Austin says the event was built on the belief that tech events should be accessible in a broad sense: closed captioning, speakers who don't all look alike, and being findable in unexpected places, like Austin's own chance encounters. Austin says conferences like KubeCon or re:Invent are marketing and sales events where learning is a side effect, and Austin was inspired by devopsdays to make something that felt like that but online. Twitch made it easy for people to jump in.

        The Keynote

        Ian's idea had been a weekly Zoom call for Kubernetes folks who missed the contributor summit turning into an informal meetup in Animal Crossing, not a whole conference. Austin thought bigger, and Ian then got picked for the keynote. Ian had planned a script on the in-game tablet with reactions and costume changes, which fell apart immediately once Ian tried to talk on Zoom, drive slides and use in-game reactions at once. What survived was the tone: warm, positive and welcoming. Ian's keynote was about DevOps and security working together with empathy, and Ian says the reactions and claps from the audience felt more real and supportive than talking into the void on Zoom. Matty adds that the speakers sat in the same Zoom all day, muted, which replicated some connection, and says other organizers should not dismiss this event.

        Dave describes it as "the best of the Internet," reminiscent of an old forum, and coordinated real-life and in-game outfits, only to learn the video wouldn't be shared. Dave and Matty both recall long pauses while hunting for the right reaction, and Matty got stuck on one reaction for three minutes without noticing.

        Why the Community Worked

        Kat says it felt like magic and doesn't know how to recreate it, since virtual expo halls with corporate-branded avatars aren't as intimate as visiting someone's island, where they spent hours and millions of bells on a little movie theater. Kat calls it "the most pure thing I have seen in tech in a long time." A participant who refers to "Katy and I" says much of it was "me getting out of the way": they gave people a space and watch parties came from attendees. Ian pushes back that Austin and Katy did a lot of active work to make the space inclusive, wholesome and accessible, and compares it to open source communities that act with care. Katy says setting that tone attracted the right people and led volunteers to offer to moderate.

        Mia says being genuine worked, but it may not be replicable, and that the tone set from the first talk mattered since Twitch chat can get mean. Mia has played Animal Crossing since 2004, and says friends and family watched and finally understood the job. Austin says about 100,000 viewers in total, including many people who weren't Animal Crossing fans or in tech, and at one point the stream ranked third in concurrent Animal Crossing streams. Ian notes a lot of security folks attended who had never been exposed to DevOps culture.

        Preparation and Speaker Camaraderie

        Kat, intimidated by the speaker list, bought Animal Crossing and wrote the talk two weeks ahead, unusual for someone who usually wings it, and showed up wearing a million-bell crown, having got lucky on the turnip market. Jacquie decided to submit after a team discussion, then got pulled into a three-day hackathon building a videogame, and proposed a talk about building a videogame from inside another videogame. Jacquie says the overlap between DevOps and Animal Crossing is building a space that empowers and welcomes people. Adrienne asked what in Animal Crossing could make awesome parallels and set out to write a CFP too creative to turn down. Katy says there were no bad CFPs and they kept adding talks, and Austin says they did add two more.

        Matty notes that Austin set up a Discord before the event, which gave speakers a way to trade slides and titles, something like a speaker dinner. Aaron started with about 60 minutes of content for a 25-minute slot, and calls single track conferences powerful because speakers who see the talks before theirs can tie themselves together and say "go listen to Katy talk about that." Matty says the event changed Matty's mind about prerecorded talks. A participant who works on a Twitch channel for a company's developer content argues that Twitch viewers expect a messier, live experience and that virtual conferences work better if you treat them like you're flying out.

        Next Time

        Austin wants another big one next year, with shorter Animal Crossing sessions over the summer and possibly other games, joking about Fortnite. Austin says 90 percent of what people saw was done the week before, and that Austin learned enough of a vector illustration tool to make rounded curves, which Matty calls the best DevOps metaphor: learning from doing. Keep an eye on desertedislanddevops.com. Matty plugs Irreverent DevOps Party Games, a live-streamed DevOps game show on Twitch on Tuesday, June 2nd at 8:00 PM Central, with Kat as an early contestant.

        Conference content
        • Videos of talks on YouTube
        • Deserted Island DevOps Postmortem
        • Shoutout to Tori Chu who both spoke and did the amazing in-game artwork and swag!
        • Recaps
          • Deserted Island DevOps Wrapup (FireHydrant Blog)
          • Deserted Island DevOps Recap (Blameless)
          • Press coverage
            • RedMonk article
            • VentureBeat article
            • TechCrunch article
            • Vice article
            • TechRepublic article
            • 1 hr 10 min
            • Security Chaos Engineering with Aaron Rinehart

              Matty and Jessica Kerr talk with Aaron Rinehart, CTO and co-founder of Verica, about applying chaos engineering to security. Aaron co-founded Verica with Casey Rosenthal and was last on the show a few years earlier, talking about taking an internal enterprise project to open source. Aaron is writing an O'Reilly book on security chaos engineering with Kelly Shortridge, and a chapter of a new O'Reilly book on chaos engineering covers the security case. The cold open is Matty's line "We're adults, but we're all kids at heart."

              How Security Chaos Engineering Started

              At UnitedHealth Group, Aaron was chief security architect and helped lead the DevOps transformation. The company hired its first SRE, who described chaos engineering, proactively breaking parts of a system, and it blew Aaron's mind, because Aaron had never seen the system and its security as separate things. The team decided that control validation made sense: you build security measures into a system with a design in mind, and need a way of continuously verifying they work as intended. Aaron was also frustrated as chief security architect that a data architect and a solutions architect would bring different diagrams of the same system, and wanted a way that wasn't subjective "to ask the computer a question." Does the firewall fire when this condition occurs? Does configuration management catch these misconfigurations?

              Aaron's short definition is a proactive methodology for understanding an inherent failure within a system before it manifests as pain, for customers or for engineers. Jessica says the key is forming a hypothesis about how the system works in some non-optimal condition and asking the real system.

              What Chaos Monkey Was For

              Aaron says Chaos Monkey began during Netflix's move from DVDs in the mail to streaming, in 2008, when AMIs were disappearing in AWS and causing outages. Netflix had no chief architect to mandate anything, so it designed services to be resilient, and Chaos Monkey would pseudorandomly take down one during business hours. Aaron says that puts a well-defined problem in front of an engineer, and "when you put a well-defined problem in front of an engineer, they solve it." Jessica says it turns "works on my machine" into a reproducible test.

              Matty adds that chaos is about testing a hypothesis that everything will be fine, not one where everything goes to hell, and quotes Netflix's line about running it in the middle of a business day in a carefully monitored environment with engineers standing by. Everybody knows the experiment is happening, and when things look squirrelly it's done, so if the key business metric heads south, pull the plug.

              What a Security Experiment Looks Like

              Aaron says most security experiments focus on accidents and mistakes, the low-hanging fruit: a weak password, ports open that shouldn't be, too much access. With 680 accounts and 200 services with conflicting IAM policies it's easy to miss a misconfiguration, so they proactively introduce these mistakes to build confidence that the tools catch them. Jessica's summary: engineers aren't perfect, so stop asking them to be, notice when they're not, and let them learn.

              Aaron's example is the open source tool ChaoSlingr, written at UnitedHealth Group, whose original name was a poop-themed joke that kept a side project fun. It had three functions, a generator, a slinger and a tracker, written in Python on AWS Lambda, with opt-in and opt-out tags. The main experiment opened an unauthorized port in AWS security groups, on the assumption that the firewall would immediately block it. They found the firewall detected it only about 60 percent of the time, due to configuration drift between their non-commercial and commercial AWS environments. The cloud-native configuration management tool, which they weren't paying extra for, caught it every time. Both tools sent log data to the security operations center, but the operators couldn't tell which AWS account and instance the alert came from, which could take hours to work out, so they added metadata to the alerts. Aaron says "Nobody's freaking out" and they learned all of it without customer pain.

              Aaron's boss, the CIO, said the tool keeps the incident team sharp by testing the tools, people, skills and runbooks. Matty adds that business hours are the best time to have an outage, since everyone is available, and practice makes incident response normal. Jessica compares it to unit tests, which give you privacy on your own computer, while chaos tests give you privacy within the company.

              Logging, Root Cause and Who Security Is

              Aaron says there is no software security logging anywhere, and that log events must be written by a software engineer and need to make sense to a human. Aaron also says "Root cause is a fallacy," and Jessica adds there are many necessary conditions, any of which could be called the root cause. Aaron says security is always an engineering problem, and that the book's audience is about 70-30 security people, trying to bring them toward the software engineering and SRE communities.

              Matty says security has come from two directions, business risk and controls in the 90s and engineering now, and that zero trust replaces the fence with locking your door. Aaron wants security in the value chain, and says DevOps helped. At UnitedHealth Group, Aaron taught over a thousand security people to write Python, not to make them engineers but to build empathy, and Aaron figures 15 to 20 percent wrote interesting scripts. Matty says ops and security are like a corporate lawyer, known only when something goes wrong, yet security and reliability are aspects of quality. Aaron also says security spending takes about 30 percent of project cost in unregulated environments and 40 in regulated ones, and mapping an experiment to the control it verifies gives "free compliance." Matty says audits are often theater and an automated trail is better than "a bunch of information that a human being typed in," and Jessica's version is "That is not an audit trail. That's a blame trail."

              What Makes It Hard

              Aaron says security chaos engineering is only about three and a half years old, there are few open source tools, and ChaoSlingr is somewhat deprecated since Aaron left UnitedHealth Group. Others are writing their own Python and bash scripts to inject failures, mostly for cloud and container security experiments. Aaron adds that one of the better tools is a Java one from a person in Berlin that hadn't been open sourced.

              Aaron is @aaronrinehart on Twitter, and there's a chance to win a printed copy of the O'Reilly book through the show notes.

              • Last time Aaron was on ADO
              • ChaoSlinger
              • Enter to win a free copy of the upcoming Security Chaos Engineering O'Reilly book
              • 55 min
              • Helm Community with Matt Farina, Karen Chu, and Matt Butcher

                Bridget talks about Helm and its community with three people who work on it: Matt Farina, who works on Kubernetes and cloud-native at Samsung SDS and has done open source for more than 15 years, Karen Chu, a community program manager on Microsoft Azure's Cloud Native Upstream team, and Matt Butcher, an engineer at Microsoft. Karen and Matt Butcher both joined Microsoft through an acquisition of the company the transcript spells "Daeus," and Bridget works on the same Microsoft team as Karen. The transcript identifies the two Matts as Butcher and Farina, and this summary does the same.

                Where Helm Came From

                Matt Butcher's one-phrase description is that Helm is the package manager for Kubernetes, and a chart is a package Helm can install into a cluster. The idea came out of a hackathon project at the company with Karen and a couple of other engineers, from wanting the package-manager experience for people starting out on Kubernetes. It's now "I don't know what, 1.8 million downloads a month or something like that."

                Karen didn't expect it to get this big, since the company was still scrappy, and recalls helping debut Helm at the first KubeCon with a turnkey booth and socks, not long after the hackathon. Matt Butcher says the booth was about 6 feet by 8 feet and KubeCon had a couple hundred people.

                The Charts Repository

                Matt Farina joined about two and a half years earlier. Farina co-chairs SIG Apps, which oversaw Helm when it was a Kubernetes subproject, and started contributing to the charts repository, a community-curated collection of packages such as MySQL and MariaDB, accessible out of the box in Helm 2. Matt Butcher says they expected 12 to 30 charts at most, and then the pull requests piled up. Farina saw a painful manual process, learned from an amazing maintainer whose reviews everyone trusted, and automated it: linting charts and pull requests, then testing installs in live clusters. That became a standalone tool, Chart Testing, now available to people who self-host repositories.

                Matt Butcher describes four phases: the Wild West, design patterns and a best-practices document written by Farina and maintainers from Bitnami, automation, and tooling general enough for anyone's chart repository. Farina adds that charts grew from simple replacements to logic and design patterns, like specifying a URL to expose an application and having everything around it created.

                From Scrappy to CNCF

                Karen organized the first Helm Summit shortly after the acquisition and before Helm joined the CNCF, and calls it scrappy and modest. Afterward, with CNCF support, logistics like the schedule and CFP were offloaded, so they could keep the conference technical, not salesy or flashy. Farina calls the first Helm Summit one of the favorite conferences Farina has been to, with different track styles, roundtable time and an intimate setting, in contrast to CNCF conferences with well over 10,000 people. Karen says they deliberately didn't scale the second one to 1,000.

                The Issue Queue

                Bridget asks how to connect with a community that wants many things. Matt Butcher says the early community was small and eager, so they could pair on Zoom or Slack with people having problems. Now it feels like "the Time to Make the Doughnuts commercial," and early issue interactions are almost robotic, asking people to fill out the template. Butcher says the biggest emotional challenge is treating each issue as filed by a person with needs, "often filing the issue out of frustration, because they don't file issues when they're happy with something."

                Farina says the way to keep focus is to say what Helm does and doesn't do: if someone has another idea, suggest a Helm plugin, or wrapping Helm, and Helm will list related projects, since it's "not a junk drawer." Generic Kubernetes questions also land in the Helm queue. They assign someone each week to triage, and Farina's best answer is better documentation and pointing people to it. Bridget agrees and mentions having sent a couple of doc updates.

                CNCF Benefits and Virtual Connection

                Karen says CNCF brings webinars, conference support, and project pavilions, such as a Helm-dedicated booth at the last KubeCon in San Diego, a neutral space for a project built by many companies. Matt Butcher says the pavilion and Helm Summits are the opposite of the issue queue, with time set aside to talk to people as people. Farina says some people just came by to say thank you, which you don't see in issue queues, and that Karen organized maintainers, more than 20 from more than 10 companies, to staff the pavilion.

                Bridget proposes a drop-in "Helm happy hour," distinct from the maintainer call, to replicate serendipity. Karen says it would remind people why they work on this. Farina says virtual conferences make speakers isolated, while a happy hour would be two-way, since "a community is a lot of people, not some people who know lecturing others."

                Phippy and Friends

                Bridget asks which Phippy and Friends character each identifies with. Karen, who Butcher says was the brainchild behind the visuals, picks Zee, who is full of questions. Matt Butcher picks the pods that carry things around until they die, and says Captain Kube came from the owl, a favorite animal of Butcher's daughter, and that Phippy was the answer to the daughter asking what Butcher does. Farina picks Phippy, for a secret PHP past and because giraffes are a favorite of one of Farina's daughters, and Butcher and Farina first worked together doing Drupal. Bridget picks Goldie, tied to the University of Minnesota's Goldie Gopher mascot and the Go community.

                Where to Go

                Matt Butcher points to helm.sh and its accessibility push for documentation, where you can learn it and re-express it in other languages. Farina points to hub.helm.sh for charts, and Karen to the Helm Twitter account and the helm.sh blog.

                • Helm.sh
                • Celebrating Helm's CNCF Graduation
                • CNCF announces Helm graduation
                • An Introduction to Helm - Matt Farina & Josh Dolitsky
                • Phippy and Friends
                • Keynote: Phippy Goes to the Zoo: A Kubernetes Story - Matt Butcher & Karen Chu - KubeCon North America 2018
                • Seven Hard Truths About Open Source Community
                • Helm chart testing
                • Helm hub
                • Helm twitter
                • Helm logo art credit: @flynnduism

                  Banner image: @bridgetkromhout

                  40 min
                • What's the Deal with AWS Billing...? with Corey Quinn and Pete Cheslock

                  Matty and Jessica Kerr talk with Corey Quinn and Pete Cheslock of the Duckbill Group about AWS bills, during the early pandemic. Matty introduces them as cloud economists, and Matty says the show will be informative and maybe hilarious. The transcript's speaker labels for the two guests are scrambled in places, so this summary uses the phrase a guest instead of pinning most stories on one of them. The cold open is one guest on how the title came about: "Do you pay money for me to be a cloud economist? They said, yes, we do. I said, yeah, I am a cloud economist."

                  What a Cloud Economist Does

                  A guest explains that the title was made up because they're two words nobody can define, then found out other people use it, including someone with a PhD in cloud economics, which led to a choice between owning up and teaming up. What the Duckbill Group does, in a guest's words, is look at companies' AWS bills, "because those tend to be the big ones," and help them become smaller and less terrifying. A guest's example of a typical surprise is an EMR cluster that fails to start but doesn't turn off the old one, and with no idempotence check spawns a new one every run, so "you're not building the cloud for what you use, rather for what you forget to turn off."

                  Matty asks why a guest who had been a cloud consultant joined Duckbill. One guest says that while consulting for a couple of years the lesson was to know only a little more than your first customer, and that the two kept saying they should do something together. The move came when projects finished early and the company was looking for people, and they slid into the CEO's DMs, a message that sat unseen for a while because of a Tweetbot bug with group DMs. As the economy got questionable, reducing spend seemed more important, since it can be the difference between laying off engineers and turning off servers nobody remembered.

                  What the Customer Says Versus the Pain

                  A guest says that what customers say and what actually hurts aren't aligned. Someone in finance sees a bill that looks like a phone number, the concern passes through about five levels of corporate telephone, and the real pain is that it's too difficult to figure out what the cost drivers are and allocate them: "understanding, optimizing, and predicting it." Now, with a recession-style pandemic event, customers who say "we're here to save money" mean it.

                  Jessica notes that DevOps was supposed to give people feedback loops and cloud took that away for engineers who can't see the bill. A guest says data centers had the problem too, buried in multi-year cycles, and that you can still do financial hijinks in the cloud. Another guest says the number of pages in a large bill can be in the hundreds, and mentions a bug found in AWS data transfer pricing where it's cheaper to transfer data between us-east-1 and us-east-2 than between availability zones. Both guests tell stories about tiny charges, one of a 22-cent charge running for years and the other of spending about 5 hours to delete a 2-cent Glacier vault, to which Matty replies that they have that 22-cent charge too.

                  Tagging and Enforcement

                  Asked about misconceptions, one guest says to tag your cloud usage, thinking about how your company makes money, since "current you is going to have the CFO roll over to you one day and say, what is our cost of goods sold?" The harder part is assuming users won't follow the policy and enforcing it, and the guest's approach is to delete untagged resources, which the guest calls the scorched earth approach, with the gentler alternative being permissions and security rules, "but no one understands IAM." A guest describes customers fixated on the wrong thing, like a company building tooling to cut its dev environment spend when development was 3% of the bill, and says an unbiased third party helps by avoiding internal narratives about the bill. Another points to Terraform plugins that estimate cost, and worries about Kubernetes, where containers make it unclear what's underneath and how to tag it.

                  Matty adds that a dollar amount without context means nothing, such as a $500 a month button. A guest adds that attributing cost to teams or users also fails without context, as when accounting asks who Jenkins is, and a data science user costs a king's ransom because that's what they do.

                  First Steps and Over-Specing

                  For the one thing to do first, a guest says to turn on the AWS billing reports, which aren't on by default and deliver data to S3, and since some take weeks to produce useful insights, do it on day one. Someone should own the Amazon bill, since "someone should have a number on their head in some way." Matty notes the old pitch was to move from CapEx to OpEx, and a guest says most people are still over-specing, picking an instance, hearing "it's slow," doubling it with no metrics, and never going back. One guest tells of quietly moving developers' unused workloads to a smaller instance type and no one noticing.

                  Asked if Lambda will fix this, a guest says Lambda solves a different problem, and jokes that many AWS blog posts conclude with "fix it your damn self" with a Lambda function.

                  Worst Names in AWS

                  Matty asks for the worst-named AWS thing. One guest says Snowball, and another says the many Systems Manager services and invents Systems Manager Cost Manager in the moment, then finds AWS Cost Categories was announced that day. A guest thinks any service starting with the word Simple sends the message that it's easy. For the worst name in cloud, a guest says Azure DevOps, because a hiring manager junk-piled a resume for listing it as if it were a skill. Jessica says it should have been called Arrested DevOps. Matty adds that you can't buy DevOps, "but I sure as hell can sell it to you."

                  Pandemic Bills

                  The pandemic brought new work. Customers made multi-year commitments assuming spend would rise forever, and now ask how to deal with commitments they may not meet. Traffic is skyrocketing for some and falling for others, but bills don't fall as much, because people "misunderstood auto-scaling to mean it only ever scales up." A guest adds that billing systems run on at least an 8-hour consistency model, so you don't learn what an expensive change cost until later. The guests close by saying that companies still paying retail prices should negotiate, since no one really pays retail, and Jessica's line on elasticity is that "the definition of elastic is not that it stretches, it's that it snaps back after it stretches, sometimes with lawsuits."

                  55 min
                • WebAssembly, Krustlet, and the Future

                  Bridget talks with Taylor Thomas, an engineer on Microsoft Azure's Deis Labs team who works on containers and Kubernetes, and Brian Ketelsen, a Cloud Developer Advocate at Microsoft who leads a group contributing to upstream open source and has been active in the Go community as a GopherCon organizer and author of Go in Action. Both work on Krustlet, a project released the week before. Bridget picked the guests by looking at the project's commit history. The cold open is Brian on Rust: "It's moving the direction of the foot gun so that it's not pointed at you."

                  What WebAssembly Is

                  Taylor says WebAssembly is a compiled language invented for the web, running in a complete sandbox, with modules importable into JavaScript in a browser. WASI, the WebAssembly System Interface from the Mozilla Foundation, lets WebAssembly modules run anywhere and not just in a browser. Brian says it's a binary format that executes anywhere there's an interpreter, so the same file works on a Linux AMD server and an ARM32 Raspberry Pi, the "write once, run anywhere promise Java brought us so long ago and we all laughed at."

                  On security, Brian says execution is entirely sandboxed and memory has to be explicitly exposed in or pulled out, so "unless there's a bug in the implementation of your WebAssembly host, it's completely secure." Taylor adds you can only do what you explicitly expose to the runtime.

                  What Krustlet Does

                  Brian says Krustlet is an implementation of the kubelet specification, written in Rust, that executes WebAssembly instead of Docker containers. It takes a pod spec, fetches the WebAssembly file and runs it, with no containers involved, so workloads are more secure, though the Kubernetes installation is only as secure as you made it. Taylor says you don't have to worry about AppArmor, SELinux or a Linux file system in a Wasm pod, which cuts the attack surface and lets platform builders worry less about Linux permissions.

                  Brian says WebAssembly on the server is a new concept for most people, and the cost angle is big, since the same files run on a 96-core ARM processor that is more energy efficient than an AMD processor. Taylor says everything runs as a thread in one process, so idle work parks its thread, unlike containerd's shim, a separate process per container, which helps on edge devices with a gig of RAM. Brian adds that WebAssembly is designed to be interpreted as a stream, so you don't load the whole file to start and the memory footprint is lower. Brian doesn't expect it to stop Internet of Things vendors from making insecure configurations.

                  Two Runtimes and an Evil Capability

                  Krustlet has two execution runtimes. WASI currently has no networking, so pods can't listen on a port. The other is waSCC, WebAssembly Secure Capabilities, created by folks at Capital One, which uses RPC between host and module and so can do network calls. Capabilities, such as logging, key-value storage and databases, are configured once on the host, and modules just request a key-value store. Brian says you can hot-swap a capability, such as Consul to Redis, without the actors knowing.

                  Brian describes building "95% of an evil thing" that afternoon, a capability called Shell that lets the sandboxed module make an RPC call to the host and run a shell command as the host process, such as ls or rm -rf. It shows capabilities are "just as smart as you make them." Taylor finds it terrifying but a demonstration of flexibility, and notes WASI is very new, with only two or three languages having strong support. Brian says this is a least common denominator interface, so you won't get every feature of each provider, similar to the Service Mesh Interface spec. WASI's specification work is done through the Bytecode Alliance.

                  Is It Ready?

                  Taylor says the repo has construction-sign warnings that it's not ready for production, since a provider lacks networking and init containers and volumes are missing, though basic pods run. They want people trying it for feedback. Outside Krustlet, Taylor says most major websites use WebAssembly in the browser, and gives Autodesk's web AutoCAD as an example, while server-side WASI has few production examples. Brian names edge providers Cloudflare and Fastly, which let you upload WebAssembly to run on the edge, and says Brian's own website runs on WebAssembly in Cloudflare.

                  Ways to Get Involved

                  Taylor suggests a demo with one Krustlet running the Wasm provider and another running the WASI provider, doing HTTP-triggered job processing, or serving a low-traffic API or web page from a Wasm module. Brian wants to turn a Go-based Raspberry Pi controller for a barbecue pit, which turns a fan on and off, into a WebAssembly module with a smaller memory footprint. Brian adds that the Krustlet process can run on anything from a Linux server to a micro:bit, extending a cluster across the globe, and Bridget reminds listeners the devices should be theirs. Krustlet is a virtual kubelet in the sense that it tells the control plane it's a kubelet, and differs in being written in Rust and not Go.

                  Why Rust

                  Taylor says the reasons are that Rust has some of the best WASI support and most of the related projects are in Rust, and that its compile-time safety guarantees around how long data lives prevented bugs. Brian says there's an opportunity cost, since Rust is harder to read and start with, and it took more than a month, closer to two, before Brian wrote useful code, leaning on Taylor. Once past the learning curve, you appreciate how much the compiler does. Taylor disagrees that Go is easier to follow as projects grow, and likes Rust's match blocks and error handling. Taylor, a core maintainer of Helm, says updating Kubernetes libraries in Go has been "an absolute nightmare," and calls Cargo "an absolute gem," with conditional compilation through features, making the dependency story "infinitely better than Go."

                  Brian streams live coding of Krustlet on Twitch, and Taylor jokes it's good for anyone with imposter syndrome. Taylor invites contributors with experience on EKS, GKE, DigitalOcean, IBM and other clouds, and stresses it's not meant to be a Microsoft-focused product.

                  • WebAssembly

                  • WASI

                  • Bytecode Alliance

                  • Krustlet

                  • Kubernetes Rust Kubelet on GitHub

                  • Introducing Krustlet, the WebAssembly Kubelet by Matt Fisher

                  • Kubernetes and waSCC by Brian Ketelsen

                  • waSCC capability that will make your ops folks cry - GitHub link to evil project (don't install this)

                  • Kubernetes: A Rusty Friendship - Taylor Thomas

                  • WebAssembly meets Kubernetes with Krustlet by Ralph Squillace

                  • 43 min
                  • Secure by Design
                    Secure By Design

                    Guests Dan Bergh Johnsson, Daniel Deogun, and Daniel Sawano join host Jessica Kerr to discuss their book Secure by Design.

                    Daniel: “There’s a lot of good designs which come naturally to us as programmers but which has the interesting side effect that they also prevent security-related bugs.”

                    Domain Primitives

                    The panel discusses domain primitives as an example of coding practices that naturally provide security through good design.

                    Dan Bergh: “It’s a good starting point to understand that using domain-driven design not only makes your code more expressive, solves more domain problems. Even though these designs were not crafted to address security to start with, they’ve also had that as a side effect.”

                    Jessica: “I love that what you’re recommending in this part is to think harder about what you do want in the system, express that in the code, and suddenly a bunch of things that you don’t want in the system just aren’t.”

                    Testing

                    The panel talks about the ways in which testing contributes to secure design.

                    Daniel Sawano: “It tends to be so much easier and more robust if you start defining your own domain types.”

                    Immutability

                    The panel discusses the benefits of immutability.

                    Dan Berg: “It’s possible to...configure and mutate them until they are kind of safe-ish.”

                    Jessica: “Kind of safe-ish?”
                    Dan Berg: “Well, we are on a DevOps podcast.”

                    Logging

                    The panel talks about the security implications of logging practices.

                    Daniel Deogan: “One thing that’s very important is that if you log input directly into your logs, it becomes an attack surface for second-order injection attacks.”

                    Dan Bergh: “It’s a perfect launchpad for doing a really, really hard attack inside your system.”

                    Daniel Deogan: “The common mistake that many developers do is that they more or less dump inputs blindly.”

                    Jessica: “We have this illusion that logging is simple, but it isn’t.”

                    Cloud Thinking

                    The panel discusses the chapter on cloud thinking.

                    Dan Bergh: “In a way, we’re instructing the system to become more intelligent.”

                    Symmathesy!

                    The book is available online in its entirety.

                    56 min
                  • Whose Transformation is it Anyway? with Andrew Clay Shafer

                    Matty and Jessica Kerr talk with Andrew Clay Shafer about transformation, recorded in March 2020 as the pandemic forced sudden change on everyone. Matty notes that the three had a 45-minute green room conversation before recording that nobody will get to hear. Andrew says "transformation" is vague and ill-defined, and like Agile, it's common enough that people think they know what it means, but "everyone has a different agenda." What Andrew cares about is how to get social-technical systems of humans and computers to do things in a way that gives the humans better experiences and performance. The cold open is Andrew's "There's always DevOps in the banana pants."

                    Forced Digital Transformation

                    Jessica notes that working from home during a pandemic, as opposed to voluntarily working remotely, makes everyone more dependent on the technical parts of the system for social interaction. Andrew says many organizations are being forced to live a digital reality they didn't want, which forces the question, though it's unclear how it will play out. Digital transformation as a tagline is not new. Andrew found an IDC projection of over $7 trillion invested in digital transformation initiatives from 2020 to 2023, made before the pandemic, and that most people consider the vast majority of those initiatives failures, with failure rates between 66% and 90%. Andrew adds that with no clear success criteria and a culture that doesn't allow failure, things that are failures by any measure are often called successes.

                    Pareto Inefficient Nash Equilibrium

                    Andrew explains the framing Andrew has used in talks for about seven years, Pareto inefficient Nash equilibrium. It describes a situation where changing certain things would leave no one worse off and at least one party better off, but no player would change their strategy unilaterally. What makes DevOps hard is that no one group can do it unilaterally, since all the players need to change behavior at once. Jessica notes that teams can act collectively but we lack a unit of collective action larger than that.

                    Andrew says it's slightly worse than that, because with the wall of confusion one group is incentivized to change things and so introduce instability, and the other is incentivized to maintain stability, sometimes with compensation tied to it. Jessica adds that those concerns are inherently intertwined, and separating them into units of responsibility creates conflict.

                    Systems Thinking and Identity

                    Andrew says most people don't think in systems. Dominant management thinking focuses on metrics, KPIs and OKRs and pushes them in the desired direction, while systems thinking argues you get more return by removing the resistance to those things happening. Those barriers get institutionalized over time. Andrew also says people attach identity to their task, particularly in the Western world "where our last names are literally connected to our vocation," so when told to do something different, they hear "you are erasing my identity" and produce an immune response. Jessica's version: "I am a Java programmer."

                    Double-Loop Learning

                    Andrew explains single-loop learning as picking a metric, taking action and checking the metric, while double-loop asks "why are these our KPIs? Should these be our KPIs?" It gives an individual or an organization the ability to update the model when the data doesn't match. Jessica says the hypothesis is most useful when it supports or invalidates a model, and the model is what's valuable. Jessica offers the example of a code change that makes something faster but also overheats a laptop.

                    Matty asks what drives leaders to think this way. Andrew says dominant management theory is an artifact of the Industrial Revolution, and applying Taylorism-driven metrics to knowledge work doesn't give good results. Andrew points to small communities pushing things forward, including Agile, DevOps, Cynefin, Wardley Mapping and resilience engineering, and to observability and chaos engineering as subgenres. Jessica says double-loop learning is needed at every level, not just among leaders. Matty notes that new ideas get retrofitted to old metrics.

                    Andrew says organizational learning is only possible if the central nervous system, back to leaders, has a good sensor network connected to line workers. If leaders don't have that mindset, they can actively prevent the rest of the organization from having it. Jessica adds it's more than changing how you think: you have to want to continually change it.

                    Doing Versus Learning

                    Andrew describes Senge's two domains, action and doing, and enduring change and ideation, both pathological in excess. The quip is "I can't learn anything, I'm too busy getting things done," attributed to the least productive person in the world, and the other danger is being disconnected from the doing. Andrew wants a balance and cycle between them, and says learning has its own exhaustion, using chess, guitar and learning Arabic as examples. Jessica adds a motorcycle riding book's advice: devote 5% of your brain to observing while riding, and in a race put 100% on going fast.

                    Matty ties it to the moment: there are times to stretch your learning and times to just do the work, and right now people don't have the bandwidth to improve everything. Andrew says "Some people are born to transformation and some people have transformation thrust upon them," and Jessica adds that right now everyone is the latter. Matty's advice is to give yourself, colleagues, family and your boss some flexibility. Andrew's final hope is that humans have "the resilience and adaptive capacity to build something better when this is over." Jessica's advice: eat more chocolate. Matty adds: wash your hands.

                    • Pareto Inefficient Nash Equilibrium
                    • Systems Thinking
                    • Double-loop Learning
                    • 38 min

                    About Arrested DevOps

                    From the publisher's feed

                    Arrested DevOps is the podcast that helps you achieve understanding, develop good practices, and operate your team and organization for maximum DevOps awesomeness.