Arrested DevOps

Arrested DevOps

By Matt Stratton, Trevor Hess, Jessica Kerr, and Bridget KromhoutTechnologyTech News
Download on the App Store

Arrested DevOps episodes

  • DevOps Isn’t a Department with Jeremy Duvall

    Matty talks with Jeremy Duvall, founder of Seven Factor Software, a software engineering consultancy in Atlanta, about why DevOps keeps ending up as a department and what to do instead. Jeremy cut teeth at Danger, which built the T-Mobile Sidekick, then worked at Microsoft, and has been doing DevOps for a long time. The episode is part of the show's tenth-year look back, and Matty promises the anniversary episode in a month. The cold open is Jeremy on the Agile and DevOps values: "focus on your people and stop worrying about the stupid shit you use to get things done."

    How DevOps Became a Department

    Jeremy's first exposure was a DevOps Days in Atlanta where John Willis spoke on burnout, which showed that DevOps wasn't just about playing with Jenkins. Before that, Jeremy says, nobody thought you needed a department to deploy Jenkins servers: two people on a tools team built the machinery the rest of engineering used to ship. The real contribution of the movement was refocusing on the idea that developers are humans. Then, as with big data and digital transformation, big companies took the ideas, put them in a department with pay scales, and turned infrastructure teams into DevOps engineers doing the same things. Jeremy sees platform engineering as "the next logical evolution" of what DevOps should have been.

    Matty adds that early DevOps deliberately avoided being prescriptive, with no manifesto, just the CAMS acronym of Culture, Automation, Measurement and Sharing from John Willis and Damon Edwards at the first US DevOpsDays in Mountain View, to which Jez Humble later added Lean, for CALMS. People still have to do work, and automation is where vendors make money, so DevOps drifted toward meaning Jenkins and Puppet. Jeremy compares it to Agile: a manifesto that gave rise to SAFe, which Jeremy calls garbage, and also to good things like Kanban and Lean. Matty adds that "SAFe is how to have Agile but let project managers still have a job."

    Measurement and Incentives

    Jeremy says the business community got its claws into both movements with Scrum, velocity metrics and somebody-has-to-get-fired accountability, which produces walled gardens, bureaucracy and a pathological culture instead of a generative one. Matty says measurement doesn't equal Taylorism, contrasts Nicole Forsgren's work on developer productivity at GitHub with McKinsey's, and names Freakonomics as the most important DevOps book: "Go learn about incentives and then you will understand DevOps."

    Jeremy says what engineers do is hard to measure, so organizations reach for features shipped or hours against velocity points, and "A measurement ceases to be a good measurement when it becomes a target." Jeremy points to the developer productivity work of Abhinav and DX, which focuses on happiness, and cites Antifragile for thinking about teams that survive someone leaving or breaking an arm. Jeremy says Google, Amazon and Facebook get this right, with engineers empowered to deploy and build their own platforms, while Jeremy's retail clients are stuck in a 1990s CEO mindset, with Nordstrom as an example of a real turnaround. Matty adds that Courtney, who led that transformation at Nordstrom, has talked about internal platforms for the enterprise, linked below.

    The Frozen Middle

    Matty describes hearing at Chef that about 80 percent of CIOs said DevOps transformation was critical but only about 15 percent of enterprises had a plan, and says the C-suite can think big while middle managers whose job is executing one way of measuring will fight a radical change. At PagerDuty, banks had whole teams of incident managers, and they weren't bad people: "I would probably have done the same damn thing."

    What Success Looks Like

    Jeremy says most of the Fortune 500 now know DevOps is a thing, and successful teams have trust to make decisions and a paved golden road to production, with a platform or DevOps team, whatever it's called, building the frameworks while software engineering teams cross-skilled in things like Terraform get their apps to production without interference. The cloud means developers no longer wait two weeks for a Tomcat server. To avoid repeating the mistake, Jeremy says to avoid commoditization: "everything that's wrong with software engineering today is commoditization and value engineering." Jeremy builds teams with overlapping skills, "I don't want specialists in 2023," and says platform teams are software engineers, "They're not infrastructure people anymore." Matty wraps up with the 2014 episode How to Eff Up DevOps, linked below.

    • John Willis’s talk at DevOpsDays Atlanta 2016 on Burnout
    • https://platformengineering.org/talks-library/internal-platform-enterprise-courtney-kissler
    • ADO - How to Eff Up Devops with Pete Cheslock, Nathen Harvey, and Randi Harper
    • 31 min
    • Runtime Analysis with Brian Kelly

      Matty talks with Brian Kelly of AppMap, a runtime analysis company, about what runtime analysis is and where it fits between static analysis and production observability. Brian is originally from Ireland, has lived in the Boston area for over 20 years, and came to AppMap from distributed systems, SaaS and a cybersecurity company. The guest's employer sells the category under discussion, which the guest says plainly when it comes up, and AppMap also appears in the links below. The cold open is Brian: "We won't say the Log4j word."

      Where Runtime Analysis Fits

      Matty's guess at the definition is "analyzing during runtime," and Brian confirms it. The difference from the APM and observability tools the industry is used to is where in the workflow it happens: on the developer's laptop or in CI, before deployment. Static analysis has become a commodity for classes of problems like vulnerable dependencies, and Brian says a team that isn't using it is delinquent. But there are problems developers assume they can only catch by sending code to production and watching it with an observability tool. Brian says "I hate to say shift left, so I'm going to say shift closer."

      The cost of finding out late, Brian says, is context switching: a developer who has moved on to another pull request hears days later that a change is a problem, and the issue tends to land on a Jira backlog, where an SRE compensates by adding CPUs and RAM, which Brian ties to crazy upside-down hosting costs. Tests and manual QA already generate runtime data that many teams ignore. Matty adds that the point is getting 80 percent: fewer defects reach production, so the remaining ones might actually get fixed, and Matty compares the backlog to Homer Simpson balancing the garbage. Brian says Dependabot and static analysis didn't end pen testing or dynamic testing, and each of them is a signal, but teams over-rely on observability tools and crank up APM spend.

      What It Finds

      Brian's best-known example is the N+1 query from an ORM such as Hibernate or ActiveRecord, where a static analyzer sees that an ORM is in use but not the queries generated at runtime or how many identical ones are issued while paginating. Others are dependency injection and dynamic library loading: a static tool sees a config file and can't see what gets loaded. Runtime analysis watches where "the infinite gets constrained down to the finite," and can show that seven or eight listed libraries are never used and whether the ones in use are used in a vulnerable way. Matty connects this to the Sysdig report finding of loaded but unused JavaScript packages in the Cloud Native Security episode.

      What It Looks Like in a Workflow

      Brian says tools in this class have to be automated, fast and fit the existing workflow: in VS Code, IntelliJ or PyCharm, pressing the run button instruments the application, watching bytecode in Java's case. That sounds like an APM, but it collects more specific data, between observability and a profiler, building a dataset of which functions, queries and libraries run. Analysis comes in automated form, like flagging an N+1 query, and human form. Brian's example: a unit test triggers six occurrences of an N+1 query that's a tiny fraction of CPU time, but the developer knows production data would make that six million. It's also definitive: "this really happened," not a fuzzy prediction.

      Brian expects the concerns to be noise, since static analyzers started out with many false positives, and expectation of magic, but runtime analysis only observes, and is bound by how long the tests take, which developers run anyway.

      The OWASP Top 10

      Brian points to the OWASP Top 10 as an illustration: around 2010 SQL injection was at the top, and then tools and frameworks commoditized the fixes. Brian recalls a 33,000-page automated pen test report for a big bank where prepared statements fixed about 27,000 pages. What replaced it are harder problems like broken authorization and authentication, cryptography and secrets handling, which static analysis mostly can't detect. A secret might pass through a framework that logs it, and a runtime analyzer can find the code paths.

      Getting Started and AI

      Matty mentions Adam Jacob's approach for Habitat, installing the gnarliest enterprise software to prove it, and says that's not the way to start. Brian's answer splits on whether you have automated tests. With tests, put runtime analysis in the editor and in CI. Without tests, instrument locally, on a laptop or a UAT environment, and look at the reports, or "personal observability," which is safe because nobody needs tests to run in production. Matty: "monitoring is simply testing with a time dimension." Brian says you should still write tests.

      On AI, Brian says most AI code reviewer apps just look at the same static diff, while runtime analysis data is a new dataset. In experiments, adding runtime data to an LLM prompt stripped out noise and produced answers like how to mitigate a known N+1 query. For anti-patterns, Brian can't think of one on the spot, and says the risk is overreach: a tool that becomes "a mosquito in your eardrum," as static analyzers did with their 79 vulnerabilities until Dependabot started opening pull requests. Brian ties the category to the DevOps hangover of overspending on observability.

      • OWASP Top 10
      • Stripe: The developer coefficient (quantifies the cost of bad code to companies to be $59B annually)
      • Facebook: FAUSTA: Scaling Dynamic Analysis with Traffic Generation (how runtime analysis was used at WhatsApp to catch design flaws before they reached production)
      • Dragan Stepanović - Async code reviews are choking your company’s throughput (from LAS 2022, a talk which highlights the systemic problems with developers trying to do manual code reviews of large PRs)
      • AppMap, the runtime analysis company which Brian works for
      • Cloud Native Security with Michael Isbitski ADO Episode
      • 39 min
      • Complexity with Michael Stahnke

        Matty talks with Michael Stahnke about whether the systems we run need to be as complicated as they've become. Michael has spent 13 or 14 years on and off the DevOps circuit, was VP of engineering at CircleCI, and now works at Flox, an 18-person company building tooling aimed at removing complexity. Matty notes the episode isn't meant as a pitch, and Michael says caring about the problem is why Michael works there. The cold open is Michael: "the original problem was I couldn't get my developer environments unified, and therefore I ended up with Kubernetes. What the fuck?"

        Are We Better Off?

        Michael's question is whether operational availability, debugging and troubleshooting are better than in 2004 or 2005, given that "what we keep doing is inventing new problems and then inventing new solutions." At CircleCI the availability struggles came from "doing really complicated shit," not bad engineers. Both recall 2003: Michael was writing software to replace spreadsheets and manual administration with SSH and for loops, managing thousands of servers, while Matty was moving from Exchange 5.5 to Exchange 2000 after an acquisition, and rebuilding dev servers at Allstate from an answer file on a floppy and a ProLiant CD while sitting in a cold data center with a book. Matty says you could hold the whole stack in your head, and asks whether microservices and distribution have left anyone better off. Michael says unequivocally yes in many scenarios and no in many others, and the first thing to understand is whether you actually have the scaling problems or just think the tools are cool. "Keeping one thing online is easier than keeping 25 things online."

        Build for the Problem You Have

        Matty recalls a Rails app being "up in 3 days, down in 3 months," and Cars.com handling Super Bowl traffic by renting servers for 48 hours instead of re-architecting, since the spike happened at the same time every year. Michael says to find users and product market fit before investing in a service mesh and service discovery: "don't underestimate the power of rsync and cron," and a colo, Linode or DigitalOcean node can go a long way. The test is whether you're delivering the value, and time spent working out whether a Kubernetes minor version will change an ingress controller API is something no customer pays for. Matty adds that "you" is doing a lot of work, from a 12-person startup to JPMorgan Chase, and tells the story of a retailer's ops team insisting on site stability until management said the job was selling things.

        How Containers Led to Kubernetes

        Michael walks back the chain: a developer environment needed to be consistent, so "we're going to package up your laptop, we're going to pass it around until it gets to production," which is what a container is. That led to a scheduler, service discovery, a service mesh, and businesses scanning containers for vulnerabilities, all layered on top instead of asking why the decision was made. If everyone developed in the same reproducible environment, Michael says, containers might not be needed, and then neither would the rest. Michael misses typing service start and strace, where a problem takes ten seconds to find, instead of launching a debug pod. Kubernetes "was originally designed to solve Google-scale problems. Unless you are Google, you do not have Google-scale problems," and Michael's image is that not everyone needs Everest climbing gear to cross the street. Matty: "What's the best container scheduler? The one you don't need."

        Organizations and Data

        Matty describes booth conversations at an AWS Summit where Kubernetes people said data was another team's job. Matty is clear that's an organizational pattern and not a failing of the engineers, and says a platform engineer should care because data is part of a platform. Matty used to accept that DevOps means never saying "that's not my job," and no longer does: "not my circus, not my monkeys" is fine, and if your job is limiting, "maybe you need a bigger circus."

        Michael adds that platform teams exist partly because the complexity was put inside the company, and asks whether they could have a smaller mandate. DevOps ideas of shared empathy and pain were good, Michael says, but operations expertise atrophied and developers reinvented tools, so a problem solved in 1995 gets rebuilt in 2018. "There's always a 27-year-old willing to redo everything you've already learned." Michael counts 150 AWS services, half competing with each other.

        What to Do About It

        Michael thinks simpler tools had a chance and missed: Docker Swarm beside Kubernetes, and a Rust tool like Docker Compose that runs plain processes without containers. Michael's wish is for a generation of tools that abstract the good patterns without all the complexity, solving the 80 percent case. Practical examples from Michael's company: a website behind a CDN and cache that could run on a Raspberry Pi with a cell modem and cost $5 a month instead of $600, and a monolith, with Knuth's line about premature optimization and "I hope that's a problem we have," since scaling problems mean users. One team with one microservice is a success; most places end up with more services than developers, and then Backstage to keep track. Michael adds that "only in software is legacy a bad word": a legacy system made the money, so don't be mad at it.

        Inside a large organization, Michael suggests shortening a workflow from 12 steps to 10, automating the repetitive debug steps, showing a decision maker two workflows and asking what you lose, and learning that "you're still a technical decision influencer." Matty recalls learning from a colleague's resume at Chase that treasury services processed $1.5 trillion in wires a day, information that sat on the business unit's intranet home page. Michael describes a Caterpillar division where every transaction ran through six servers, about $6 million a day, which made an $80,000 software upgrade an easy ask. Michael suggests reading an S-1's risk section to learn what matters to a business, and both lament that value stream mapping is discussed less than it used to be, with Matty blaming Steve Pereira no longer going to DevOpsDays.

        Michael's closing ask: understand the outcomes at the other end, even approximately, so that "my complexity is built because it achieves this goal," and send in stories of solving a problem with a simple solution like installing an RPM and hitting start. Sometimes, Michael notes, you can't do that 300,000 times, and then you actually do have the problems the tools were designed for.

        47 min
      • The Database Calls are Coming from Inside the House with Grant Fritchey

        Matty talks with Grant Fritchey, who last appeared on the show in December 2014 in The Database: The Elephant in the Room, about what has and hasn't changed for data and DevOps in the almost ten years since. Grant is still at Redgate Software and has added PostgreSQL to the skill set, and Matty now works at Aiven, a data platform, which comes up as context. Both describe themselves as having "storied careers, which is a nice way of saying we are getting old." The cold open is that line from Matty.

        Nobody Has to Explain DevOps Anymore

        Matty frames the problem: code is close to immutable and easy to roll back, while data is living, and orchestration tools work well until someone asks about the data and it becomes somebody else's problem. Grant says that ten years ago every data person needed an explanation of what DevOps even was, and now "you don't have to talk about or explain what DevOps is anymore," which saves half an hour of every talk. Grant uses Donovan Brown's ordering of people, process and products.

        Matty asks whether the organizational divide has improved, based on conversations at an AWS Summit booth where infrastructure people treated databases and Kafka as another team's job. Grant says it's getting way better, with a split between companies that build DevOps in from the start and incorporate data management, and companies that add DevOps later. Grant also notices that database administrator is a term going away: people who do backups, availability and query tuning now call themselves data engineers. Matty compares it to sysadmins becoming DevOps engineers with a pay bump, and both say to change the title if it helps, since the work matters more.

        Still Struggling

        Grant admits being "a bit of a negative Nelly" here: "We are still struggling." The obstacle is persistence, since you can't toss the database and start over, and deployments have to happen without taking the server down for three hours. Grant thinks data management people haven't explained well enough to developers what they need, and developers often look at the database and see something scruffy. The joke that data work is like sweeping up behind a parade with elephants in it is, Grant says, how it sometimes feels, though the job is fun. Teams that have been bitten by data problems either embrace automation or avoid the topic altogether.

        Databases Are Code

        For the platform engineer who runs Kubernetes and leaves data to a data team, Grant's message is "Databases are code," and Matty adds "just big code." It's big, so it can't move fast and self-provisioned development databases need third-party tools or empty databases, but it can be treated like the rest of the code, with the special part being persistence. Grant says the biggest misconception is that the database can't be automated. It takes a bit more discipline than automating development because the data must be kept, but it's fully automatable, and wildly successful organizations automate their data management and deployments.

        Communities and History

        Grant finds the PostgreSQL community as welcoming as the SQL Server community is most of the time, and notices a split in Postgres between committers oriented around academics, theory and science, and everyone else who uses it and wants help. Grant says the same split is familiar from DevOps, where "Peggy does DevOps" doesn't make a DevOps company. Matty draws on the history-of-databases talk with Kat Cosgrove, from 1970s Ingres through Postgres, and says knowing why things are the way they are helps. Matty recounts that the relational model's creator disliked rows, columns and tables and wanted something more mathematical, like tuples.

        Grant says the old foundations persist because they work, "rebar inside of concrete," which the Romans used. Grant is building a LoRa and IoT project on Azure and still uses a relational store, since the data volume is small. Specialized databases such as Cosmos DB are like a specialty tool for taking antennas off a radio, and you still need a hammer and nails.

        What's Exciting

        Grant is excited by open source being everywhere: AWS, Azure and GCP all include or support it, and the divide between the commercial and open source camps is shrinking, though paid software will remain, as with Query Store in Postgres on Azure. Grant also expects people to try to put AI into production tooling.

        Grant's advice for growing a career is to "assume automation from the start," then decide whether you enjoy the why of things, such as query tuning and design, or the how, such as automation in Azure and AWS, Kubernetes and containers. The local job market matters too: in Tulsa, Oklahoma, SQL Server dominates, while on the coasts Postgres is growing. Grant has submitted to PGConf Europe in Prague. Grant's last word: "Treat your database like code."

        • Arrested DevOps - The Database: The Elephant in the Room
        • Arrested DevOps - Data! Data! Data! With Francesco Tisiot
        • Arrested DevOps - The New DevOps With Adam Jacob
        • History of databases talk from Matty and Kat Cosgrove
          43 min
        • Platform Engineering goes to Flavortown with Matt Kurtiz

          Matty talks with Matt Kuritz, staff engineer and tech lead of the platform engineering team at The Farmer's Dog, a fresh dog food company, about what platform engineering looks like at a company that has been doing it for about four years. The company has grown from a few engineers to roughly 50 to 100 and has shipped over 100 million meals. Matty frames the episode as a return to platform engineering after recent episodes with Daniel Bryant and Pete Cheslock, with the joke that none of them actually do it and just want to talk about it. The cold open is Matt Kuritz, on why tools are better: "infrastructure tools and software just becoming more like real software."

          Product Teams and Platform Teams

          Matty starts from the Charity Majors post on the Honeycomb blog, which has a table contrasting platform engineers and ops engineers, including a row where SSH is a no for platform engineers. Matt Kuritz says the idea came partly from a diagram in Lean Enterprise about self-service operations building a PaaS, which Matt never wanted to build, but its goal stuck: in an ideal world there are no handoffs between developers and operations, and product teams own their software end to end. At The Farmer's Dog there are two kinds of team, product engineering and platform engineering, and "the only difference is who the customer is." Everyone cares about reliability and delivery. Matt Kuritz adds a disclaimer that if SRE or DevOps works for a company, there's no reason to drop it, and that some companies have to go deep on infrastructure.

          Matty adds that nobody can talk about platform engineering without James Governor's line about "endlessly remaking remakes of Heroku," and says the point of Heroku was getting to value fast, with abstractions where they need to be, and not feeling like Dropbox. Matty repeats that "the best tool is the one you don't need. The second best one is a SaaS." Matt Kuritz says to "do what works at the right scale and then iterate and evolve from there," naming Vercel for front-end-heavy companies and Stripe as a business function nobody wants to run. The Farmer's Dog isn't building a PaaS but assembling a platform of preferred tools with glue where it pays off, which Matt sees as a pattern others could use.

          How the Team Measures Itself

          Matt Kuritz says the team has to be aligned with product teams on the goal, and their north star is ownership. The first milestone was continuous deployment, since every commit then has one clear owner who is responsible through the pipeline to customers in production. The team reached 100 percent continuous deployment of its applications after fixing a distributed monolith, and before any golden paths. Matt is frustrated by takes that say DevOps is dead and platform engineering is the future, and recommends Charity's DevOpsDays New York talk. The work, Matt says, wasn't different from what a DevOps team might do, but the language and structure changed: platform and product, not dev and ops. Ownership is still a multi-year goal, with each application having one clear owner, and "platform doesn't own anything."

          For the product management side, Matt prefers to work hands-on with internal customers: do a real use case manually, maybe one or two more times, and only then invest in a code generator or golden path. "We don't upfront really decide anything. It has to be done at least once and put into production." The mistake teams make, Matt says, is making decisions in their own echo chamber.

          Who Decides What

          Matty's rule of thumb for what to standardize is whether something is an interface point between groups: a shared source control tool matters, the JavaScript form validation library doesn't. Matt Kuritz says platform steps in as the decider for cross-team choices, such as asynchronous messaging, where RabbitMQ, SNS, SQS, Kinesis and Kafka could all show up in one app. The team collects opinions from all teams and asks whether a tool that becomes the standard would be worth maintaining, and challenges assumptions. Matt says Kafka is powerful and flexible, with many ways to get hurt, and the team evaluated Confluent but couldn't find use cases yet. Matt describes path dependence with driving on the right side of the road, and says to be honest about migration costs. Matty adds that people who've never heard of a tool will say it's fine for their team, so the platform team has to bring the context.

          Bounded Contexts, Not Silos

          Matt Kuritz says "functional silos are unhelpful," but not having boundaries means everyone has to know everything, so the answer is bounded contexts in the domain-driven design spirit: vertically integrated teams mapped to business domains, like fulfillment and signup, that rarely need to coordinate. Jess Kerr's blog post on better coordination or better software supports the idea: build technical components that let people work asynchronously instead of perfecting the handoffs.

          On the technical side, the team chose a monorepo, Node and one language, and not introducing another until necessary. The monorepo exists to practice trunk-based development and continuous deployment, and works only if there's one version, which is head. Upgrading Fastify across dozens of apps is one atomic commit that triggers the tests of dependents. The team has no Backstage or internal developer portal, since a code owners file in a repo organized by business domain is the directory, and Git is what developers already know. "What's our platform? It's a monorepo and a code generator. That's our platform."

          Start and Stop

          Matt Kuritz's thing to start is setting a long-term vision you can never fully achieve, an idea from Toyota Kata: for The Farmer's Dog it's that every dog lives its longest life. Then break it into 1 to 3 year challenges and near-term target conditions, like moving a 10-person team to three deploys a day. The thing to stop is copycatting, or cargo culting, which Matt ties to Feynman's Cargo Cult Science and to Lean being copied from Toyota by its artifacts like Kanban cards. In platform engineering it's launching Kubernetes and Backstage because other companies did. "Think about your problems, solve those problems from first principles, don't copy."

          • Arrested DevOps - DevOps With Better Marketing with Pete Cheslock
          • Arrested DevOps - Platform Engineering with Daniel Bryant
          • Arrested DevOps - Platforms with Kelsey Hightower and Andrew Clay Shafer
          • Lean Enterprise
          • The Future of Ops Is Platform Engineering
          • Charity’s talk from devopsdays NYC
          • Jess Kerr’s blog that Matt mentioned
          • Cargo Cult Science
          • 47 min
          • What's Up With Open Terraform?

            Matty talks with Ohad Maislish, co-founder and CEO of env0, which manages Terraform, Pulumi and CloudFormation, and Cory O'Daniel, CEO and co-founder of Massdriver, a visual environment for cloud infrastructure that depends heavily on Terraform, about the community fork then called Open Terraform. Both guests run companies with a stake in the outcome, which Matty says up front while insisting it's not a criticism of their motives. Matty's own history is with Chef and Pulumi, and Matty says there is no dog in this hunt. The episode's URL and title predate a rebrand announced partway through. The cold open is Matty: "definitely GitHub is just effed right now."

            What Happened

            Ohad's account, as a participant: on August 10th HashiCorp changed the licenses of several projects, including Terraform, and a group of vendors and individuals wrote a manifesto asking that Terraform stay open source forever. When HashiCorp kept its decision, which Ohad calls "totally legit," the group started working on a fork. Cory adds that some CNCF projects began pulling away from HashiCorp tools, and that Massdriver wants to bet on a tool that the open source community keeps investing in, the "new lingua franca of infrastructure as code."

            Matty asks why a license change matters to an ordinary engineer. Matty's concern is vagueness: depending on how a lawyer reads the license, distributing a project that uses Terraform could be a violation, and you can't always know who is competitive with HashiCorp. Ohad says the license was followed by a binding blog post of clarifications on August 21st and changes to the Terraform Registry terms on August 24th, which shows the community didn't understand the implications. Every fundraising round, Ohad says, means sending the license to investors, and a BSL license with a binding blog post and more updates is harder to explain than a clear alternative. Ohad contrasts it with SSPL, which exists to stop others from reselling a vendor's product as a service, and calls BSL flexible and, in Ohad's opinion, still vague.

            Matty adds that enterprises already fight compliance just to use open source at all, and a complicated license makes that harder. Cory says Massdriver was affected and then wasn't, since it competes with Waypoint, which was carved out of the FAQ a week later, and that dynamism is what worries people. Cory gives examples of transitive dependencies: a CNCF project that uses Consul, projects pulling out Vagrant, and Jaeger considering removing a Go plugin, which is still MPL. "It is very much email us to figure out if you owe us money," Cory says, compared with how other projects handled BUSL changes.

            OpenTofu

            Asked what to call it, Ohad says OpenTF, and Cory breaks the news that will go public the day after recording: the Linux Foundation recommended a rebrand because TF could be confused with Terraform, so the project becomes OpenTofu and the binary is tofu. Ohad and Matty make the fork and tofu jokes. The project now belongs to the Linux Foundation, the natural step toward CNCF, and Ohad says it is meant to be a drop-in replacement: no code changes, just a different binary and registry. Gruntwork, the creators of Terragrunt and Terratest, and Harness are among the supporters, and Gruntwork has released a sneak peek of state encryption, which Matty recalls as a Pulumi differentiator since a Terraform state file can hold secrets in plain text.

            Matty asks what makes the fork more than a burst of pledges, recalling a maintainer maxim: "as a maintainer, no is temporary. Yes is forever." Ohad points out that Terraform core was open source but stopped accepting community pull requests about two years ago, while OpenTofu has a five-member steering committee under Linux Foundation and CNCF guidelines, and has already deferred a community change to a later version as a new capability. Ohad says RFCs and voting rules are a work in progress. Cory hopes that teams replacing HashiCorp tooling will contribute and speed up a project that "really has been bogged down for the past couple of years."

            The Registry

            For modules and providers, Cory says the first registry will proxy requests to GitHub and resolve module names to repositories, and Ohad adds that provider naming conventions map to GitHub URLs. So publishers won't have to submit anything. Cory sees an opportunity to support OCI-compliant registries later.

            How to Help

            Ohad suggests starring the repo, joining the community Slack, filing and voting on issues, and sending pull requests, and sees an opening for large vendors and cloud providers to shape the project under CNCF. Cory suggests putting OpenTofu into CI pipelines, for example with Terratest, once the alpha registry is out, to surface edge cases in loading providers and modules. Cory's closing point is that the group is "a consortium of competitors" that get along well because they love the language. Matty promises to file a first pull request to nominate Cory's Instagram-famous Australian Labradoodle, Ziggy O'Doodle, as mascot.

            • https://www.instagram.com/ziggy.odoodle/?hl=en
            • https://twitter.com/opentofuorg
            • https://github.com/opentofu
            • https://linkedin.com/company/opentofuorg
            • https://github.com/opentofu/opentofu
            • DevOps World is back for 2023, and you won't want to miss out on this one-of-a-kind event! This year's program is packed with exclusive insights, immersive workshops, and unparalleled networking opportunities taking place across multiple cities in the US, UK, and Asia. Elevate your DevOps game and register using the following links: NYC area, Chicago, Silicon Valley, Singapore, and London.

              39 min
            • The New DevOps with Adam Jacob

              Matty talks with Adam Jacob, CEO of System Initiative and, in a previous life, CTO of Chef and the person who wrote Chef originally, about what DevOps got right, where it fell short, and what a second wave of tooling might look like. Adam counts DevOps from John Allspaw and Paul Hammond's 2009 talk at Velocity, and describes being on the "loves being a systems administrator" side of the field. The episode is sponsored by Uffizzi, and System Initiative is Adam's own company. The cold open is Adam: "It didn't matter how good you were at operations if the application didn't run, and it didn't matter what your application did if there was no infrastructure to run it on."

              What DevOps Got Right

              Adam starts with what went well by describing the year 2000: operations was separate from IT, siloed, and ran on long planning cycles, with developers requesting gear that took six to eight months to arrive and capacity planning a big deal. The organizations worked, Adam says, a lot like how people describe platform engineering today: operations stitches it together and builds systems so developers don't think about infrastructure, "and never the twain shall meet." The top achievement of the DevOps movement was recognizing a single continuum of work, and the pre-DevOps world was "objectively worse" in day-to-day work, tooling and what the systems could do.

              Matty adds a story from leaving Chef: a customer engineer at a large Chicago financial organization apologized for not getting much done, and Matty, who'd watched them in the thick of it for years, said they had come a long way. Adam agrees and says the flip side is real, with teams wanting 100 deploys a day with no friction and still mostly making pipelines and hoping it worked out. Following the DevOps Handbook roughly as written, Adam says, leads to mediocrity, deploying once every six months, which is better than before but rarely great.

              The Platform Engineering Smell

              Adam says the platform engineering rebrand comes from the feeling that DevOps failed because outcomes didn't arrive, and the proposed fix "does smell a lot like what life was like in 2001": "let ops be ops, let engineers be engineers" with some software in between. Matty puts it as accepting that the silos can't be removed, "So let's just make sure the silos are better." Adam's answer is that the people who succeeded collaborated better, and that the tooling was never designed for collaboration.

              Adam's history of automation runs from 1990s compute clusters and college labs, through ratios of 10 to 1, then 100 to 1, and up to 10,000 to 1 as EC2 and Facebook apps arrived, and each stage had to make up an answer for the world of the time. DevOps then "ossified the shape of the world into that shape." The Flickr talk shows it: a portal for deploying, feature flags, dark launches, configuration management, capacity planning and metrics, with deploys started from Subversion tags, so they "literally built a platform roughly the way that we describe it." Every piece of the stack has been rebuilt ten times in ten years without an appreciable change in outcomes, and Adam insists it isn't any one tool's fault.

              Tools and Culture

              Matty credits Adam with "tools influence culture, culture influences tools," which ends up in Matty's decks. Adam says regretting "the idea that it's about culture and not about tools was wrong from the jump," because culture is what you do, and a version of culture in DevOps was like being a lapsed Catholic: believing the right things hard enough. An organization will not become more collaborative without tooling that forces it, since individual people won't do it alone. "Tooling is culture because it's literally what we do all day."

              Matty ties it to the book Switch, where a manufacturing machine that kept injuring hands was redesigned so that turning it on required both hands: make the right way the easy way. Matty's Asana fight over a defined process is the same thing.

              Factories, Soccer and Collaboration

              Adam says DevOps inherited factory metaphors from lean, but software is closer to a professional sports team or an orchestra: highly motivated specialists unified on one objective and making tiny decisions together in real time. Adam's picture is a soccer team retrofitted with pipelines, review by other people and no coaching during a Scrum window. The result was process and tooling that sucked out the creativity and collaboration, where people end up "working near each other at best." The most important insight of DevOps was that what separates great teams from okay ones is "the rate of collaboration," and nobody designed tooling for it.

              Matty asks how much of this is outside the control of the people who can change it, comparing it to Agile needing changes in finance, sales and marketing. Adam answers that change is hard at scale but not insurmountable, and that the LivingSocial story of sales having already sold a removed experiment is a collaboration failure. At System Initiative, which Adam describes as "everything I believe at 12," one Miro board runs from the pitch deck down to an individual story, and Adam reads the whole strategy to everyone every Monday so engineers know why and for whom they're working.

              Scale and the Enterprise

              Matty asks about JPMorgan Chase scale rather than a 15-person company. Adam says large organizations can be convinced with data that DevOps is better, but they buy tools instead of building them, unlike Google and Facebook, whose tools are bespoke. The vendor's temptation is to adjust the tool to the customer's culture for a bigger check, which is "a tomorrow problem for tomorrow people," and Adam recalls a Cloud Foundry deployment that collapsed under its own integrations. The way out, Adam says, is to build a tool that teaches the right way, because enterprises reliably buy tools that promise better outcomes: "let the tool change them, not the other way around."

              Matty pushes back that the cynic sees churn, and asks whether the tool will get used by the people in the middle. Adam says it will if it's good, pointing to Salesforce as sticky because "it's actually kind of good," and admits not being sure this is the only answer. What Adam is certain of is that "the status quo is not a good enough answer," since a new Terraform-like or Ansible-like tool will have roughly the value of the old ones. The thing not yet tried is changing the shape of the system: why a pipeline in the middle, why only one layer of the application, and "what even is an application?"

              Simulations Instead of Code

              Matty asks how System Initiative reasons about this. Adam starts from automation, which has followed one line of thought since Mark Burgess's computer immunization paper in the 1990s, and argues that code is a poor medium for collaboration and that feedback loops of 20 minutes or more, with an opaque Terraform plan, are too slow. Looking at adjacent fields, Adam lands on digital twins: Formula 1 teams simulate every variable of the car so only promising changes are tried at the track. System Initiative is an attempt to build a high-fidelity simulation of infrastructure, removing glue code and building the workflow into the system, and Adam says more people need to pull on other threads too.

              Adam's closing hope is that people become willing to change the shape of the system and "throw some of the babies out with the bathwater." Matty adds that the past was the best available with the information at the time, and that maybe it was what got us able to learn these things.

              Links to Resources Mentioned
              • 10+ Deploys Per Day: Dev and Ops Cooperation at Flickr
              • *The DevOps Handbook
              • The System Initiative
              • Switch: How to Change Things When Change Is Hard
              • DevOps World is back for 2023, and you won't want to miss out on this one-of-a-kind event! This year's program is packed with exclusive insights, immersive workshops, and unparalleled networking opportunities taking place across multiple cities in the US, UK, and Asia. Elevate your DevOps game and register using the following links: NYC area, Chicago, Silicon Valley, Singapore, and London.

                59 min
              • Purposeful Personal Brand with Cassandra Faris

                Matty talks with Cassandra Faris, who runs community for KubeCampus, a free Kubernetes training product, about building a personal brand on purpose. Cassandra began in tech as a recruiter, has a philosophy degree, and went on to manage communities around SaltStack, a financial project and now people learning Kubernetes. The cold open is Cassandra: "the more you do that and the more you reach out and help other people as well, the stronger your brand is, the stronger your reputation is, the more success you have."

                You Already Have One

                Cassandra's practical case is that a brand helps a career: speaking invitations come from what people know of Cassandra's work, and for jobs "the resume is kind of a formality at this point." Even an individual contributor needs some online presence so people can see what they're about. Matty adds that you have a brand whether you want one or not, since it's what people think of when they think of you, such as the SRE everyone asks about replicas, or the colleague who is difficult to work with. Matty brings up the Mad Men line, "if you don't like what people are saying about you, change the conversation," and says that we have some responsibility to take ownership of it.

                From Accidental to Purposeful

                Cassandra's own brand started by accident. As a recruiter told not to be the one who doesn't know Java from JavaScript, Cassandra went to Agile and cloud meetups, then followed everyone tweeting the #CodeMash hashtag, listened for a while, and started to participate, posting job search tips and skills advice. Developers turned out to welcome a recruiter as long as Cassandra was authentic and not working an angle, and invitations to panels, talks and conference organizing followed. Cassandra realized that being deliberate about what was posted could work in Cassandra's favor, hence the talk on a purposeful personal brand.

                Matty's version: at Chef, Matty was known for infrastructure as code, and at PagerDuty, where incident response was the focus, nobody knew Matty for that despite years of doing it as a job. So for a year Matty took any stage to talk about incident response, blameless postmortems and learning from incidents. Matty insists that this wasn't inventing a persona, it was deciding what to be known for. When Matty's boss later asked for the same reputation in the ITIL and IT service management world, Matty's answer was that "you can't just decide to have authenticity." It takes time, and you have to be able to back it up. Cassandra says every role change means updating, not rebranding: the human skills of building relationships with a community's key leaders are the same whether the community is open source contributors or Microsoft MVPs.

                Your Brand Inside a Big Company

                Matty notes that almost nobody stays at one big employer for 20 years, so external brand matters even in roles that aren't public, and asks how it works inside an organization like Target or Visa. Cassandra says to choose opportunities that match who you are: not competitive, cutthroat cultures if you aren't competitive, and plenty of room for creativity if you need it. Matty adds that the tension is credit: talking about what the team did as well as what you did, since a story that is all about I doesn't show a team member, and ops roles are like a corporate lawyer, where nobody knows what you do until you don't do it. "If you don't, who's going to do it for you?" Nine times out of ten, Matty says, you think you're bragging and you aren't.

                Three Questions and Framing Accomplishments

                Cassandra suggests writing down answers to three questions, which are in the slides linked below. Who are you professionally: your top three technical specialties, three professional specialties and three contributions to the team and profession. Who are you personally: three interests and hobbies, three beliefs and values, and three parts of your social, family or community life you want to connect over. And how do you connect with people: shared interests, advice you're seeking and what you want to know about others. That gives you a bank of material. "Modesty isn't a career-enhancing trait, but that doesn't mean be an asshole." The trick is to frame accomplishments by who they helped. Cassandra's example is a KubeCampus Kubernetes Learning Day that sold out and gave 200 people an introduction to Kubernetes, which reads as "it helped all of these people" and not as bragging. Give credit to the team, and practice, even in a mirror.

                Show, Don't Tell

                Matty says the quickest way to get mocked on tech Twitter is to put thought leader in your bio, and recalls putting up a page of DevOps thought leaders a decade ago until J. Paul Reed asked to be taken off it. If you solve a problem, the story should be about the interesting way it was solved, and the expertise comes through. Matty also recommends keeping a brag book, an Evernote notebook of kind messages that started as a cure for imposter syndrome at Chef and turned out to help at performance review time and with seeing what's worth telling. Matty's DevRel team writes a monthly wrap-up for the company, framed around why it's useful to sales and others, and not as a pat on the back: "how do you brag by giving someone something useful?"

                Making Time and Choosing Platforms

                Cassandra keeps an hour blocked on the calendar every morning for social media, takes a day or two off the internet after a conference to avoid burnout, and pays attention to whether the people and things followed make things feel psychologically unsafe, since that makes content hard to produce. Platforms are shifting as Twitter usage declines: short-form video is growing, and Cassandra sees a generational split where Gen Z and younger millennials want to connect on Instagram while everyone else uses LinkedIn. The brand can stay the same while the tools change, so it helps to be on several.

                Matty suggests blogging may have a renaissance, especially documenting what you're doing or learning in the open. Matty's most visited page is a post on configuring SharePoint in a one-way trust, written to document a fix for a team, which became the answer on an MSDN forum. Matty tells of Annie, who learned an infrastructure testing tool in public on a blog until the product team said "Annie's blog is our documentation now." Beginner content is the hardest to create and the most needed, Matty says, since hundreds of thousands of people need to install Kubernetes and only a dozen need to optimize sharding at a million users per second. Teaching also helps you learn.

                Vulnerability, Asking and Paying It Forward

                Cassandra is seven months into a role at KubeCampus, is teaching the networking fundamentals lab, and plans a blog series on the learning journey. "It starts with admitting you don't know." Nobody has ever responded by doubting Cassandra knows what a cluster is, they've offered to explain it. Cassandra's tip for new brand builders is to ask a question or for resources about something you want to learn, since developers love to share knowledge.

                Matty adds that asking for help to learn is different from asking someone to do a task, and appreciation matters without genuflecting: "I'm not going to carry the bucket for you," but I'll give you a map, and sometimes you just need the bucket carried and can ask for that favor. Cassandra says helping people isn't transactional but pay-it-forward: Cassandra can't hire a former boss and mentor who is looking for a director of DevRel role, so the return on that help is paying it forward, and Cassandra passes the same knowledge on to others. Matty says to ask few favors and accept no, and that you don't need to go to the most famous expert every time, since thousands of people know Kubernetes. Cassandra's answer is meetups and conferences, where the hallway track has led to meeting an author whose Kubernetes book now helps, and a KubeCampus Slack for when learners get stuck.

                Links to Resources Mentioned
                • Keep a Brag Book
                • KubeCampus: Free Kubernetes Training
                • Purposeful Personal Branding - Nov 2018 Slides 13-15 have the 3 questions

                • *DevOps World is back for 2023, and you won't want to miss out on this one-of-a-kind event! This year's program is packed with exclusive insights, immersive workshops, and unparalleled networking opportunities taking place across multiple cities in the US, UK, and Asia. Elevate your DevOps game and register using the following links: [NYC area](https://reg.rainfocus.com/flow/cloudbees/devopsnyc/webinar3/page/landing), [Chicago](https://reg.rainfocus.com/flow/cloudbees/devopschicago/webinar3/page/landing), [Silicon Valley](https://reg.rainfocus.com/flow/cloudbees/devopssiliconv/webinar3/page/landing), [Singapore](https://reg.rainfocus.com/flow/cloudbees/devopssingapore/webinar3/page/landing), and [London](https://reg.rainfocus.com/flow/cloudbees/devopslondon/webinar3/page/landing).*
                  53 min
                • Everything's a Product with Sarah Morgan

                  Matty talks with Sarah Morgan, senior product manager at Telemetry Hub, about applying product management thinking to DevOps, SRE and internal platform work. Sarah has about ten years in product after an engineering background, and a career that goes back to the early 2000s. Matty's talk Everything's a Product, linked below, is the starting point. The cold open is Matty: "I think everything's a DevOps problem."

                  Thinking Like a Product Owner

                  Sarah says product people tend to stop at the business case and the features and forget reliability, stability and the rest of the user experience. Ownership has also gotten more siloed as systems have grown, so nobody sees the big picture. Matty recalls a product owner for the SRE team at PagerDuty and failing to find any posts or talks from that person. Sarah contrasts the two roles: supporting a SQL Server cluster meant caring about the health of the cluster and little else, while a product manager has to care about all of it, whether or not every piece is understood, because it all affects the end user.

                  Matty's line from the talk is that you won't put an NPS score on your Jenkins pipeline, "but kind of are you," since it's about users and feedback loops, and that's DevOps. Sarah suggests a health rating for every piece of the system, covering whether it does what's needed to support the people paying the bills. Sarah has mostly worked in B2B software, where users often have no choice, and says internal stakeholders should be treated the same way: don't build things so hard to understand or so locked down that colleagues can't make sense of them.

                  A PM on the DevOps Team

                  Sarah was the product manager for a DevOps team at a company whose product was a messaging gateway for an IoT platform, a role the company hadn't had before. The team, spread across Boston and Budapest, managed development environments, AWS infrastructure, federation access control, security and disaster recovery, and was about six people before Sarah made it seven. After watching for a month, Sarah saw repetitive work that could be productized and automated, and spent much of the time running interference between the team and people used to going straight to a favorite engineer for access. The team's customers were the software developers, so the job was finding their pain points and building a strategic roadmap, where ops work is usually "pipeline-style ticketing" of whatever is oldest or loudest. The team liked having a direction.

                  Roadmaps and Promises

                  Matty's talk has a slide that says "this is why roadmaps are bad, and if you have one, you should feel bad," since a roadmap can be read as a promise, and asks how to keep transparency without that. Sarah calls it the bane of every product manager's existence and handles it sneakily: only the roadmap shared with engineers is called a roadmap, and that's where dates live. Everything else goes out as vague documents with names like "H1 priorities." Sarah tells sales to say "new alerting integrations" and not whether it's PagerDuty or a webhook, and adds detail as a release date firms up.

                  Matty recalls Marty Cagan spending two days at Apartments.com, and the LivingSocial story of running an experiment, removing it, and having sales say it was already sold. Matty's point is that you can't be agile in only one part of the company, and ties it to Andrew Clay Shafer's observation that everyone wants more reliability, stability and velocity without changing anything.

                  Platforms as Products

                  Matty says platform engineering is "fundamentally providing Heroku inside your company," and that an internal platform team might think it needn't worry about customers because the CIO mandated it. But people who don't get what they need will go around it, which is "shadow IT all over again," and platforms tend to be thought of as compute and orchestration while event streaming, data pipelines and messaging get left out. Sarah says a SOC audit reveals the rogue platforms, and that you have to think about use cases: at a previous company a backend service didn't handle bulk requests, a front-end engineer looped over a couple thousand rows, and it fell over as soon as two people used it at once.

                  MVP Means Prototype

                  Matty says Marty Cagan reads MVP as minimum viable prototype, meant to help you learn, even though most people hear minimum viable product and treat it as version 1. Sarah says experiments are supposed to be fast and needn't scale, but then they ship and never get a version 2: "MVPs are not MVPs anymore." Matty compares it to critical production systems under someone's desk, including a machine in a Bank One data center that nobody could identify. Discovery, Matty adds, means asking questions and finding proxies, since people know what they want but don't always communicate it in a form you can build. Sarah says to repeat requirements back and to have conversations, because a PRD or ticket can't hold it all.

                  Product Managers as Incident Commanders

                  Matty says the two places product people shine are community conference tables, where they spend the day talking to users, and incident command. At PagerDuty a product owner asked to join the incident commander rotation, which had been engineering management. Told that a product owner wasn't an engineer, the answer was that "that's a feature, not a bug." The product owner became the first non-engineering incident commander, and by the time Matty left there were no engineers in the rotation. The reasoning: "never half-ass 2 jobs, whole-ass one job," so an incident commander who isn't a software engineer never gets pulled into fixing. The skills overlap too: prioritizing, communicating and delegating. Matty says it also changes how product folks think about reliability, because the concerns arrive firsthand and not filtered through on-call.

                  Matty tells of a sysadmin on the team who couldn't get the product owner to care about a service throwing around 25,000 false errors a minute, because the sysadmin led with the symptom, and the argument that landed was that nobody would notice a real failure. Putting the error rate on the office dashboard got it fixed in about two days. Sarah's version is selling disaster recovery to a general manager by describing the worst outage scenario in dollars, and Matty adds that "money is a lingua franca," not that it's all anyone understands.

                  Sidecar Skills and Learning

                  Matty calls business cases, writing and speaking "sidecar skills" that set engineers apart, recalling developers who dropped an idea when the CTO asked for a business case, and saying the exercise works like a rubber duck. Sarah says none of the interpersonal, interviewing, listening or public speaking skills came naturally and all were learned. Sarah also warns against underestimating colleagues' ability to learn technical details, from sales to product. Matty extends it to incident commanders, who need to understand a system well enough to know whom to call, without knowing how to rebuild the carburetor.

                  Sarah's start and stop: start thinking about who your stakeholders are, and stop building things without validating them.

                  • "Everything is a Product" - Matty’s talk
                  • “More Buzzwords Won’t Help” - Andrew Clay Shafer
                  • Telemetry Hub channel on YouTube

                  • *DevOps World is back for 2023, and you won't want to miss out on this one-of-a-kind event! This year's program is packed with exclusive insights, immersive workshops, and unparalleled networking opportunities taking place across multiple cities in the US, UK, and Asia. Elevate your DevOps game and register using the following links: [NYC area](https://reg.rainfocus.com/flow/cloudbees/devopsnyc/webinar3/page/landing), [Chicago](https://reg.rainfocus.com/flow/cloudbees/devopschicago/webinar3/page/landing), [Silicon Valley](https://reg.rainfocus.com/flow/cloudbees/devopssiliconv/webinar3/page/landing), [Singapore](https://reg.rainfocus.com/flow/cloudbees/devopssingapore/webinar3/page/landing), and [London](https://reg.rainfocus.com/flow/cloudbees/devopslondon/webinar3/page/landing).*
                    51 min
                  • Cloud Native Security with Michael Isbitski

                    Matty talks with Michael, who began as an enterprise architect at Verizon, moved into assessing application security there, spent about five years in research and advisory at Gartner focused on application security, and is now director of cybersecurity strategy at Sysdig. Sysdig sponsors the episode, and the report the conversation starts from is Sysdig's own. The cold open is Matty: "And that's how we did security in the '90s, yo."

                    What the Sysdig Report Shows

                    Michael says the report is based on anonymized customer data, not a survey, so it covers a slice of the industry: organizations that have acknowledged a security problem and use tools like Sysdig. The number of vulnerabilities is alarming, which says something about the state of open source and the hygiene of components, and scanning usually reveals a worse picture than anyone expected because of nested and transitive dependencies. One result Michael double-checked was the share of non-human identities, which dropped from 88 percent to 58 percent of identities in customers' cloud environments, a shift Michael attributes partly to organizations staffing up after the pandemic.

                    On SBOMs, Matty notes that everyone at KubeCon in LA at the end of 2021 wanted to talk about them and the report suggests the industry is mostly still talking. Michael says there are two problems: the SBOM formats aren't settled, and an SBOM has to be dynamic, since a system drifts from its design over time, and has to account for partners and suppliers as well as your own code.

                    Build Dependencies, Runtime and Noise

                    Matty points to the report's finding that fewer than 1 percent of JavaScript packages are in use at runtime, and guesses it's because build tooling is written in JavaScript. The DevOpsDays site is a static site generator with a long package.json, none of which ships in the built artifact, yet Dependabot calls it "insecure as hell." Matty then catches a mistake in that reasoning live: the site does load Bootstrap on the front end, so there is front-end JavaScript after all.

                    Michael adds that a website tends to accumulate JavaScript libraries, then marketing adds tracking and payment processing adds more. A scanner will list every dependency and known vulnerability, but it can't say whether the code is reachable at runtime, which is where Sysdig's runtime insights and "in-use exposure" come in. Without that, organizations are "flying blind" and taking a best guess at what is exploitable. Michael recalls engineering teams suppressing findings in open source libraries they don't own, which gets described as false positives, and Matty calls it normalization of deviance. In cloud native, with microservices, containers and ephemeral resources, the dashboard can be "a sea of red."

                    What Shift Left Means

                    Matty says "nuance is hard" and the shorter the phrase the more nuance it needs, and recounts writing a talk about shifting left securely out of frustration with a customer's sysadmins, on the idea that a title doesn't imply infallibility. Michael describes the traditional waterfall model with security as questionnaires and compliance, and shift left as pushing security into early design and automating tests in the IDE, at commit, in CI/CD and at runtime. Each stage produces scan results, which creates a correlation problem, and many of the problems of waterfall come back, just earlier.

                    Matty's definition: "shifting left to me is not shifting the work to the people on the left. It is actually moving that domain expertise earlier in the conversation." That's why NoOps never happened, since ops turned out to be a domain of expertise, and expecting software engineers to absorb InfoSec is unfair and a bit insulting to security people. Michael says many of the calls at Gartner were about pushing scanning onto other teams because the security team couldn't scale, and notes that dynamic scanning tools are "glorified fuzzers" that need good test automation to reach the functions, and that release decisions and pass or fail builds are still unsolved. Matty adds that monitoring is "just testing with the time dimension," so a check in pre-production and in production should be in parity, and that the old hardening sprint produced a note from the security team saying it was okay, which bad guys on the internet don't care about.

                    Matty says the tooling has improved: security scanners once cost around $15,000 a seat, and code instrumentation was limited by cost, so tracing ran on one of 30 servers. Both agree that doing this right means changing how product release and sales think, and not just engineering.

                    Security as a Feature, Privacy and Zero Trust

                    Matty asks whether organizations treat security as a product feature, in the sense of a secure product and not AuthN. Michael says marketing still wants security features, and that more training material exists through groups like OWASP. Privacy has brought more focus, which Michael traces to GDPR, and so has the US National Cybersecurity Strategy and SEC disclosure mandates.

                    Michael builds zero trust from least privilege: zero trust is "least privilege on steroids," assuming the environment is compromised and authorizing continuously. It includes zero trust network access, which Michael says is what people usually think of, along with BeyondCorp, a cloudified VPN, and BeyondProd. Like shift left, people cling to one actionable piece. The report says "90% of granted permissions are not used," which Matty ties to onboarding: a new hire gets cloned from a colleague's access, the same way Chef-era teams copied an existing VM to get a new Apache server, until nobody knows what's on it. At Matty's current company, new hires get almost nothing and get annoyed, which works because the culture accepts 30 to 60 days of ramp up.

                    Confessions

                    Matty's Netflix password is a variation on the domain admin password from Apartments.com almost ten years earlier, and it was never changed when someone left, because in four or five years only one person who knew it left. "You can have bad password policies if people don't quit." Michael admits to simple passwords that a spouse can remember, since a 3-year-old leaves little time for a password manager. Matty closes with Windows NT 4.0, where passwords expired after 30 days with warnings at 15 and no minimum age, so the team built a Visual Basic app that changed the password 15 times and then back again. Michael will be at a Gartner Security and Risk Management Summit and on LinkedIn, and Matty ends with Gartner's "rogue sessions," where an analyst argues against Gartner's own position.

                    • Sysdig 2023 Cloud-Native Security and Usage Report
                    • Pushing Left With Tanya Janca (ADO epsiode)
                    • Shifting Left Securely (Matt's talk)

                    • *DevOps World is back for 2023, and you won't want to miss out on this one-of-a-kind event! This year's program is packed with exclusive insights, immersive workshops, and unparalleled networking opportunities taking place across multiple cities in the US, UK, and Asia. Elevate your DevOps game and register using the following links: [NYC area](https://reg.rainfocus.com/flow/cloudbees/devopsnyc/webinar3/page/landing), [Chicago](https://reg.rainfocus.com/flow/cloudbees/devopschicago/webinar3/page/landing), [Silicon Valley](https://reg.rainfocus.com/flow/cloudbees/devopssiliconv/webinar3/page/landing), [Singapore](https://reg.rainfocus.com/flow/cloudbees/devopssingapore/webinar3/page/landing), and [London](https://reg.rainfocus.com/flow/cloudbees/devopslondon/webinar3/page/landing).*
                      56 min

                    About Arrested DevOps

                    From the publisher's feed

                    Arrested DevOps is the podcast that helps you achieve understanding, develop good practices, and operate your team and organization for maximum DevOps awesomeness.