
Sign up to save your podcasts
Or


Matty calls this the first "rogue session" of Arrested DevOps: a second look at platform engineering, after the episode with Daniel Bryant, with returning guest Pete Cheslock, who skipped that episode and comes in "unadulterated." The first half is Pete's and Matty's take on whether platform engineering is anything new. The second half is about Pete's side project of recording people pronouncing tech words. The cold open is Pete: "If charisma was an open source app or a product, it would be called Riz."
Pete's hottest take is "DevOps with better marketing." The immediate reaction was the unoriginal one: isn't that what you've been doing all along, building a platform for your team? Pete cites James Governor's tweet that "we've spent a decade rebuilding Heroku poorly." Matty traces the lineage back through PaaS, with Azure's original web and worker services model, Engine Yard and Heroku, all of which boil down to getting software running somewhere faster without worrying about the bits that run it. Matty also admits to owning and maintaining Pete Chess Bot, which lived only in Heroku, with the code never put on GitHub, until an unpaid Heroku bill got it shut down.
Matty recalls an agile transformation in 2010 at Apartments.com, where agile coaches said infrastructure was a service that gets consumed and so didn't need to be in the conversation, which is "why the DevOps movement had to happen." Matty pushed for sysadmins embedded in squads, but with too few of them, each ended up on three squads with no time for real work after the ceremonies. When something can't be dedicated to one feature team, Matty says, an abstraction layer is the answer, and that's the evolution into platform engineering in theory.
Pete has built this kind of thing three times in a decade: a knife command to spin up servers with Chef on EC2, provisioning bare metal with Chef at a DNS company, and a framework at ThreatStack where developers filled in the application's shape and committed code. The problem is a numbers problem, with about 100 developers and 3 systems people, and self-service is how to stop being the bottleneck. The missing piece at most companies, Pete says, is that nobody acts as product manager for the ops team, whose customers are the developers.
Matty says that if you offer a platform, it's a product and you have to treat it as one, referencing a talk called Everything's a Product. A platform team may think adoption is guaranteed because the CIO chose the tool, but that attitude is what produced shadow IT. Matty also says a lot of platform engineering conversation is "incredibly Kubernetes-focused" and treats the runtime as the platform, while data and observability get left out, and then the platform team ends up integrating with 15 different ways to do Kafka. Pete says picking Kubernetes because everyone does is the same as "no one gets fired buying IBM."
Pete ties it to what Pete calls "the DevOps hangover": 10 to 15 years of growth where nobody minded how much Datadog or AWS got consumed, and now someone asks about the Datadog bill and the cost of running Kubernetes, and half the team has been laid off. "I guess we should have just had product managers the whole time." Matty's cynical read is that a product owner on a platform team would be one of the first people laid off. Pete notes that product management is the least defined role Pete has seen, from mini CEO at some companies to owner of one feature at others.
Matty's worry is "the wrong way but faster," the Simpsons line about the max power way, where platform engineering is the SRE team, which was the rebranded ops team, which was the sysadmin team. A coworker once said of 12 years at the same desk and five different companies that the job never changed. Matty adds that "silos are okay" as long as domain experts work together earlier, and describes arguing about shift left and security on Twitter. Pete jokes that if platform engineering pays 24 percent more than SRE, Pete will be a platform engineer. Matty says that's good for the individual, but asks whether a team still doing tickets and requests on Kubernetes instead of vSphere is really treating it like a platform.
Pete says the step still missing is requirements gathering, which is talking to the other people in the organization about what outcome they want before writing the first YAML: "It's always a people problem." Matty adds that Backstage is a tool for building the portal, not a platform, so you can't "rub some Backstage on it." Platforms are a socio-technical system, and buying one takes buy-in from the people who do the work and the executives who fund it, with the frozen middle in between. Matty says selling OpenShift at Red Hat was hard for that reason, at half a million to $2 million.
Pete's story is a DNS company where "platform V2" became a trigger word. Pete wrote a ten-page product requirements document, renamed the project Honey Badger, and talked to every team to distill the minimum: provision a server anywhere in the world with a default operating system and a Chef role. It worked only with cross-team buy-in and budget. Matty also plugs a DevOpsDays Chicago talk on organizational politics.
Pete's side project started with a long-standing idea to record friends pronouncing tech words in a game show format, and a Slack teasing about how Pete says a load balancer's name, for which Pete made up an origin story. Pete now works at AppMap, a 10-person company that let Pete record the series, with weekly posts across YouTube Shorts, TikTok, Twitter and LinkedIn. The first iteration is supercuts of about 20 words, about 30 seconds each, from 31 interviewees. Pete's favorites were SQL, epoch, fsck, which has about 20 alternative pronunciations on Wikipedia, and JWT, where the docs say it's pronounced "jot" and nobody Pete recorded said it, until some had implemented it.
Matty says that's bad marketing, since nobody would search for jot. Pete says some pronunciations exist to help a listener type the word, like /etc and /lib, and that "all pronunciations are valid," since the goal was never to make fun of people. Pete wants more non-native English speakers for a season two, and a form for words and volunteers is linked below.
Matty's advice on names is to ask everyone how they want their name pronounced, even names you think you know, which is what DevOps Party Games did, and offers "which one do you prefer" as a better framing than asking which is right. The episode closes on words that still trip people up: Pete's words like through and throw while reading aloud to a child, and Matty's arugula, which Matty replaces with rocket. Pete's daughter, after a show says a word from a book they read together, looks over and says, "wow, you weren't even close on that one."
Matty talks with Sidney Miller, who has worked in talent acquisition for 24 years on the back-end infrastructure side, with a focus on candidate experience and on getting marginalized populations into tech. The conversation is about layoffs, how to protect yourself mentally while out of work, pay transparency and how to negotiate. Matty opens with a disclaimer: the stories are "so cleverly disguised" that nobody should try to work out whether they're about Matty or someone else. The cold open is that same line.
Matty describes hanging out a shingle as open to work about a year earlier and watching layoff waves roll through from the end of 2022, while the same companies kept hiring, which could be a left-hand, right-hand problem or "eventual consistency." Sidney says early layoffs hit functional roles like engineering, product and data science, and since March Sidney has seen many sales, tech talent acquisition and HR people on the market, which Sidney reads as a sign of overhiring and a course correction. Even experienced engineering managers are struggling, because the open roles are so specific that the ask is for a "purple unicorn with a T-Rex on its back shooting off machine guns."
Sidney was laid off in September of the previous year from a Series B fintech when interest rates went up, was picked up at a Series A startup after leaning on Twitter, and was hit by a second reduction after that. The common denominator across all of these is that "when someone loses their job, they lose their, they feel like their value is gone."
Matty calls this "Rule 7: don't take it personally," a phrase from a former therapist, and says it's nearly impossible because so much of identity gets tied to the place you work. Matty recalls bleeding "Chef orange" at Chef, feeling betrayed when the company did things Matty disliked, and needing two jobs to learn the distance. Now Matty loves the company, but "the company is not me," and colleagues are coworkers: "we're not a family."
On layoff decisions, Matty says the best performers sometimes get let go, since the best performer may be the most expensive, and describes an organization where people couldn't volunteer to be laid off in place of a colleague because the decision was about a function and not about performance. Matty adds that managers who had no part in the decision are the ones who have to make the calls, and that CEOs crying on YouTube about it are tone deaf.
Sidney's advice for the mental side is to "give yourself space to process" before firing off resumes that aren't ready, to build a support system, and to find a cheerleader, because people in fight or flight end up on a hamster wheel. Sidney adds that a therapist can help. Sidney also notes that even with 24 years and a history of startups that were bought for billions, leaving still felt awful because the identity was wrapped up in "I built that."
Matty describes the early 2000s, when a company acquired by GE laid the team off, the job search dragged on, and Matty took a cut of around 30 percent that took years to recover, since the next employer asks what you make now. Matty now takes recruiter calls for the information: asking for ranges, seeing what people are wanting to pay for a role, and watching people who need the work accept it. Matty suspects some companies see a chance to reset pay and get people cheaper. Sidney passes on advice from a woman CMO that it's okay to take a step back in uncertain times and cover what you need to, and that you can always keep looking while working.
Sidney says employers in California, Colorado, Connecticut, Maryland, Nevada, New York, Rhode Island and Washington are not allowed to ask what you made before, and that roles posted as remote that could be hired in those states tend to disclose a range. Matty says some ranges are games, like a posting spanning over $200,000 where nobody's paying the top number. Sidney says recruiters ask about threshold or expectation, which has nothing to do with last year's W-2.
Matty walks through Talk Pay: Lauren Voswinkel started the #TalkPay hashtag in 2015, and Paul proposed an in-person version as an open space at DevOpsDays Austin that year. Matty has helped run them at many DevOpsDays and did two at Austin in 2023, anonymous but with people free to volunteer their number. Matty says people fear that their number is too low or too high, and that in practice nobody thinks the lowest number is the person's fault: "They think your employer is bad." The links below cover Paul's write-up and the Corey Quinn and Sonia Gupta talk on salary negotiation.
Sidney's tactic is to ask what the level is for the role and what salary corresponds to that level, instead of naming a number, because "the first person to name a number loses." If the answer is higher than expected, keep a poker face and celebrate privately. Hiring managers should not hand out different salaries to people doing the same job at the same level, which Sidney calls out where it happens.
Sidney tells of a woman engineer whose offer came in $40,000 below what a same-level person in Portland, Oregon was getting. Sidney pushed to match the salary, and the engineer cried and said thank you. "It's not rocket science. It's influence."
Matty adds that most of the audience can't yell at talent acquisition, so the lever is transparency from peers. Matty describes telling a friend exactly what the offer was, and a case where a friend applied to the team without asking. Matty's advice to people who aren't cis white men is to find someone who looks like Matty to coach them on having audacity: before pushing back on a wrong title in an offer letter, Matty asked a friend to "remind me to be a white guy in tech." The recruiter's reply was that it was a mistake and the DocuSign needed regenerating. And if an offer is withdrawn because you asked for more, Matty says, "good."
Sidney's last advice is to stop telling yourself you aren't worthy of the next opportunity, to process what happened through a support system, to ask for the level and the band, and to be as confident as possible. "All they can say is no. You got a 50/50 shot." Matty says a follow-up episode is likely.
Matty talks with Francesco Tisiot, a data professional with 15 years in the field, including 12 years of consultancy across a wide range of sectors and about ten of those with big enterprises. The conversation covers how data moves through a company, what event streaming and change data capture are, why data sprawl is a technical, financial and security problem, and four directions for judging whether a data platform is robust. The cold open is Matty, after Francesco explains replaying events from Kafka: "I'm still stuck back in 1999."
Matty asks how the field has changed from the days when the DBA was the person in the closet. Francesco picked data because the work sits between computers and people: analysts, scientists and engineers translate business expectations into models and documents. "Data is data, but how you think about it, how you make sense of the data, changes every time you interact with different people."
Francesco frames data as a journey, using a company that sells shoes online. A purchase lands as a record in a transactional database. A startup then runs analytics queries against that same database to compare today's sales with yesterday's, which works at small scale. As the company grows, those queries make the transactional database suffer, so the data gets moved into an analytical database, and one piece of technology becomes a chain of them. Francesco notes the chain follows the growth of the company, including change data capture, event-driven architecture and analytical databases.
Francesco explains the old batch approach: wait for the night when the website is offline, extract everything from the transactional database, and load it into the analytical database to feed the data marts and warehouses. That still suits reporting on last month's data. It doesn't suit cases that need an immediate reaction, such as reordering inventory when ten pairs of shoes sell. There the work moves from batch to real-time or near real-time, per event. With Kafka, a streaming technology, a change in the database can be propagated to downstream systems, for example an application that compares current stock against a data scientist's prediction.
Change data capture tracks changes in the database by reading its logs, without continuously querying it. Francesco says it lets a company evolve from batch to event-driven without touching the transactional database or the application in front of it, so the business keeps running while the backend changes.
Matty asks how large organizations keep their data interactions aligned. Francesco describes enterprises where you don't know about a data mart sitting there, or who purchased a tool, because "people come and go and company remains and data remains and, you know, bills remain." With non-technical users building analytics through point and click, there is no longer a single view of reality. Francesco admits losing a ten-year battle against people exporting transactional data into Excel, and says what matters is a global view of which data assets and pipelines exist and how they connect. "If you don't have the map, you are lost."
Without the map, a company can pay a consulting firm six months later to solve an inventory problem it already solved, or let two teams solve the same problem and end up with different answers because of a small detail in a KPI definition. Francesco adds the financial and security side: GDPR aside, loose control over where data lands invites someone to dump a database export into a bucket by mistake.
Matty asks whether anyone does this well. Francesco has seen two approaches work. One is documentation, which is risky because it's an afterthought, though Francesco wonders if ChatGPT could help with parsing it. The other is a single tool for all transformations, such as Informatica, which can give column-level data lineage, but data now spans many technologies and teams. Matty says relying on documentation and training is the worst way and prefers guardrails that make the right way the easy way, and notes that Kubernetes is great until you want state.
Francesco sees a way forward through metadata. Within one database like Postgres, catalog views list tables, users and permissions. Between systems, the gap can be closed by parsing the configuration that connects them, for example the JSON payload of Kafka Connect, which says where data comes from and where it goes. Automated tooling can describe how a pipeline was built but never why, so the reasoning behind a KPI, such as using the last six months of sales instead of three, still needs documentation.
Francesco's framework has four directions, and the episode's existing links include the blog post on it.
Francesco says some of these will bite immediately and some in the long term, and looking at all four lets you compare technologies and make a 360 degree decision.
Matty finds versioning and replay hard to imagine, since even restoring batch data to a point in time was nearly impossible. Francesco describes Kafka as a log: events are written one after the other, and reading doesn't delete them, so another application can read the same messages. Add schemas, and Kafka refuses messages in the wrong format. Schema evolution lets you change the schema while existing consumers can still parse messages. In the example, a boolean for putting initials on the shoes is added, billing ignores it and the printing team uses it. If the flag later needs to be a color, the log can be kept for days or years and a new set of consumers can be fed the data generated since yesterday, five hours ago or an hour ago.
Asked about what developers and infrastructure folks get wrong about data, Francesco says the biggest misconception is that data systems are old and boring. The data world is hot, and ChatGPT and AI are data. Innovation with data depends on quality data, which is the biggest problem to solve everywhere, and "Data pipelines are cool."
Francesco is on Twitter as @ftiziot, which is consistent almost everywhere apart from LinkedIn, and the DMs are mostly open. Francesco also says the main mission in life is telling people how to properly eat Italian dishes, which for example excludes carbonara with cream, pineapple on pizza and cappuccino with lunch or dinner.
Matty talks with Daniel Bryant, who runs the DevRel team at Ambassador Labs, about platform engineering: what it means, whether it killed DevOps, and how to approach building a platform. Daniel started as a Java engineer, did software architecture and then operations, and built platforms on Mesos and Kubernetes. Matty refers back to an episode recorded in January 2016 with Kelsey Hightower and Andrew Clay Shafer that talked about platforms before the buzzword existed. The cold open is Matty's opening to a topic that might be on the tips of listeners' tongues.
Daniel's formal definition is the discipline of building toolchains, workflows and platforms to support the team going from idea to observable business value in production. Continuous delivery doesn't end when the app reaches production, since feedback from a business or operational point of view has to come back. Matty says that means Kubernetes alone isn't a platform, and cites James Governor's line that everyone is trying to build their own Heroku.
Matty asks about "DevOps is dead" marketing at KubeCon. Daniel, with 20 years in Java where "Java is dead" recurs, says a shock headline draws attention, and that engineers want to bury the previous generation to look like they're innovating, while history rhymes. Matty says the question is what you mean by DevOps: if a DevOps team is an automation team that builds infrastructure, platform engineering replaces that, but the principles have been the same throughout. Matty adds that DevOps takes research seriously, and the research behind some of that marketing was talking to three companies.
Daniel says some people say IDP for internal developer platform and others for internal developer portal, and the difference matters: Backstage is a great jumping-off point, but Daniel sees it as the UI, CLI, SDK and API on top of a platform. So Daniel asks whether people mean a platform soup to nuts or just a service catalog. Matty likes the idea of a service catalog that catches the 80 percent case and lets you do other things at a cost, and raises Charity Majors's maxim that the best tool is the one you don't need and the second best is a SaaS, asking whether there's any SaaS for this, and whether Heroku was the closest.
Daniel says Heroku and Cloud Foundry were built for web monoliths, but now the 80 percent includes machine learning apps, front ends and systems of record. Spotify talks about golden paths, such as one for machine learning apps, one for microservices and one for front ends, so there's no longer one true way and a one-size-fits-all platform hits at most half the apps. Matty says that makes it a big ask to build an internal SaaS, treated as a product.
Daniel borrows a phrase from Team Topologies, the thinnest viable platform. Large organizations like Intuit, which has a global team supporting its platform, can justify it, but for a startup of three people without product-market fit, "please don't build a platform," and something like Cloud Run or Knative with GitHub Actions is enough. Matty says the tech is the easy part, the hard part is how people communicate, and cites a line that source control is a communication tool for developers. Matty also cites Adam Jacob's point that tools influence culture and culture influences tools, a variant of Conway's Law. Matty says large platforms like OpenShift or Tanzu are hard to start small with and get decided at the CIO level, and Daniel says the cognitive load is high, while bottom-up requirements plus mid-management buy-in and then exec support works better than the old golf-course selling.
Matty says to begin with the end in mind, without solving everything, but avoiding decisions that make it impossible later. Daniel says this is the role of a platform product owner, who starts small and thinks big.
Daniel says platform teams are often drawn from infrastructure or developer experience and enablement teams, and often a pain triggers them, such as an exec discovering 10 versions of a platform and nobody who knows how to maintain version 9. Matty says this is the early adopter part of crossing the chasm, and that typical enterprises haven't arrived, recalling arguing with PagerDuty's founder that enterprises still have production support teams and calling it survivor bias. Matty worries the late majority skips the pre-work and just renames teams: tech ops became cloud ops, then DevOps, then SRE, then platform, with the same remit, and says "get your bag" to anyone getting a better title.
Daniel describes Backstage as a jumping-off point for an internal developer portal, with a service catalog, automation hooks, search, who's on call and ownership, and TechDocs for living documentation. Spotify sells extensions, and Roadie offers Backstage as a service since it's hard to install, and many people use it as a facade or for inspiration, for templates that spin up a new service with observability and security baked in and links to dashboards.
Daniel's three keys are treating the platform as a product, since you can't have good developer experience without good user experience, focusing on workflows and tool interoperability, and making it composable. The thing not to do is buy a platform. Matty adds "you can't buy DevOps, but I can sell it to you." Daniel says late adopters worry about being left behind and throw money at it, while it's better to read the many blog posts from companies that have built platforms and spend time understanding the problem space first.
Matty, coming back for a new season in 2023, talks with Casey Rosenthal and Courtney Nash of Verica about the VOID, the Verica Open Incident Database, and what its research says about incident metrics. Casey is best known for chaos engineering: Casey wrote the definition with a team at Netflix, started the conferences and the community broadcast, wrote the book on it, and is now CEO of Verica. Courtney is a former chair of the Velocity Conference who worked at O'Reilly, Amazon and Microsoft before taking a research role at Verica, and calls themself "the world's only internet incident librarian." The cold open is Courtney: "We've got a bunch of data that have allowed us to bust a few myths."
Courtney started collecting incident reports while doing product research on Kubernetes and Kafka, looking for non-marketing information on how they fail in the wild. Having gathered about 1,000, people said thank you and "we have more of that," and the database and its metadata emerged in late 2020, with a first report in 2021 and a second more recently. The database takes a broad view of incidents: postmortems, status page updates, tweets and media articles, anything where someone talked about a website or service falling over. It also collects metadata such as how long the organization said the incident lasted, severity and methodology. Courtney wants data, not abstraction, to address myths about software incidents.
The VOID focuses on availability incidents since security breach databases already exist. Courtney says the DevOps mentality of sharing failure is a long way off in security. The analogy Courtney uses is the US airline industry in the 1990s, where the push to share incidents came from pilots, not regulation, and they set competitive concerns aside. Courtney and Casey both worry about regulation of software, with Courtney citing the call to regulate software after the Southwest Airlines meltdown, which was an organizational and cultural problem, and Casey calling it a personal nightmare. Casey says the VOID can help redefine the value of availability work, be it incident analysis, response or chaos engineering, and that availability and security are two sides of the same coin for system safety.
Courtney explains that MTTR comes from physical manufacturing, where widgets wear in predictable ways with a normal distribution of repair times. Software incidents don't look like that: the duration data are heavily skewed, with a big bump under an hour or two and a long tail, so "you can't take averages of that. It's just garbage." An engineer from Google wrote an O'Reilly report that ran Monte Carlo simulations on incident durations, shortening some and taking averages, and the results were a mess. The VOID team did the same with its own data and got the same results. Courtney says people react either with "oh shit" or "oh shit, but I'm going to fight you on it." Casey adds layers: the statistics are wrong, the data coming in is garbage because methods for determining how long an incident lasts are fraught, and you'd need more data points than Google has.
Matty says the metric is used to report on the performance of people within a quarter and is about the closest thing to putting a Nagios counter on your people, and recalls incident command workshops where people said they spent their time on mean time to innocence. Courtney asks what decision you would make based on it, since either way you'd have to go look at what's happening in the system and the people operating it. Courtney says the industry's data-driven obsession overlooks data about human behavior, which is still data.
Casey says organizational researchers call software the bureaucratic profession, bought wholesale from manufacturing and scientific management. Every role, such as people manager, project manager and architect, takes responsibility for expertise away from engineers, which makes sense on an assembly line and is the wrong model for knowledge work. The business must change if it wants availability, resilience and security. Courtney's line: "Taylorism is a corporate disease that we haven't developed a vaccine for yet." Matty recalls the Upton Sinclair quote about salary and understanding, and the frozen middle, where people's jobs exist because of metrics like MTTR.
Courtney says reliability and security are now central to business concerns, as Southwest showed. Casey describes how every Netflix engineer knew how to look up stream starts per second, one metric correlated with value, and Netflix found that going from four nines to five was pointless because users' Wi-Fi and ISPs wash it out, so they invested in regional failover. Courtney says to find the core metric according to your business, and Matty notes public sector people know their mission better than private sector people know how their company makes money.
Courtney says what replaces MTTR is the recognition that no single metric captures reliability, and that software systems are sociotechnical, so you need people skilled in social science methods: interviewing, collecting stories, building narratives and finding patterns. Organizations are hiring incident analysts, and Courtney believes the ones who invest will gain a competitive advantage, though the data for that is further off. Courtney's favorite example is an organization in the CIO's office at IBM, described in a DevOps Enterprise Summit talk, where a skunkworks team did quality incident analysis, a bad incident gave the opening to try something different, and now monthly learning-from-incident reviews draw hundreds of people. Courtney also notes that the Microsoft Azure team changed its public reports from RCAs to post-incident reviews and the quality and depth improved dramatically.
Matty tells how a PagerDuty SRE quietly changed the word "root cause" in the postmortem template without asking permission, nobody objected, and suggests the homework for PagerDuty users: change the template to contributing factors. Casey "fully supports small acts of vandalism." Courtney's last finding is that the VOID data shows zero relationship between an incident's duration and its severity, and Matty and Courtney note severity is negotiable and changes during an incident. Casey says the whole model of severity is wrong at scale, since some group of users is always unable to reach your system.
Jessica Kerr talks with Roni Dover, a developer who has also worked as a product manager and describes oscillating between the two, about continuous feedback as the missing loop in DevOps. Roni is a board game fan and a self-described skeptic, and started an open source project called Digma to put the idea into practice. The cold open is Roni on code: "the tales that this code could tell, if only it could tell what happened back when it was, you know, used or abused."
Roni says development processes try to optimize for speed of deployment, cadence and time to release. If you only optimize for speed, "you're just creating a system where you're hurling features over the fence faster," 24 times a day instead of once a month, without improving the learning or the feedback. Jessica points out DevOps is supposed to keep caring after production, and Roni says the tools for that look for problems, which is reactive and generic. As a product manager Roni had tools like Google Analytics showing the impact of a decision, such as whether a navigation change increased adds to cart, and didn't feel like they were running blind, while developers had CI, CD and testing tools that say nothing about impact or performance in production. Jessica calls the missing piece observability, and adds it must be more than monitoring.
Roni describes continuous feedback as the inverse of continuous deployment: the DevOps loop takes code from source to production, while continuous feedback starts with information from production, goes through stages to work out what's relevant, and ends back in the developer's tools, including the IDE. Roni stresses it's not actually linear, since feedback also exists before you start coding. Code ownership has grown from "done when I sent it to QA" to owning tests and deployment, which Jessica compares to moving from owning a car to parenting, where you want to know how your kids did in kindergarten.
Roni's examples of what code could tell you: how heavily the code is used and whether it's a bottleneck in a high-concurrency environment, which tells you what to optimize, whether it runs in production at all, which Roni has seen surprise teams who invested three years in a feature whose code path was never reached, and how it scales with concurrency, database size or payload. It could also report runtime errors, such as whether a "should never happen" branch does. Roni adds that developers misuse logging for this and forget to check.
Roni says continuous feedback makes the organization a learning process instead of a shipping process, and stretches the definition of done. Without it you accumulate technical debt and end up "running around, putting out fires." Roni cites biases: estimation anchoring and optimism bias, and confirmation bias in tests, which codify expectations and miss things nobody thought of. Observability injects relatively objective data, and "if you know about a bias, it seldom helps you actually overcome it." Jessica adds that combinations of features, data and ordering go well beyond edge cases.
Roni says very few engineering organizations actually practice continuous feedback, for three reasons: engineers are busy and can't keep looking for trouble in logs and dashboards, not all have the expertise, and they get data, not insights, and context switching is costly. Roni says an insight would be that this is a bottleneck and why, with a way to double-click for more.
Roni started Digma, which is open source and entering beta, to tackle those three issues by codifying the tribal knowledge about how to measure latency and read time series, bringing it into the IDE so there's no context switch, and making it proactive, so the code sends "life signs" after it ships. Roni also wants to celebrate wins and not make observability all about blame. Readers can sign up at digma.ai, and mentioning the podcast gets a bump up the beta list.
Roni says OpenTelemetry was pivotal because everyone agrees on it, with vendors aligning around it and a spec that lets new open source tools make the data more useful. The libraries auto-instrument code, so getting from no telemetry to useful data is quick. Digma works as a pipeline and not an APM, ingesting OpenTelemetry data, scanning the code to correlate data to locations, and, if you add the commit ID via an environment variable in CI, relating insights to the code change that precipitated them, which Roni describes as a matryoshka design. Roni would like shorter loops, so adding a trace gives feedback in testing, CI and staging quickly. Jessica relays a friend's wish for something that says there's an N+1 query right here, and Roni says that's exactly Digma's point.
Roni says a developer wants control, not a 2 AM call three days after a push, and had written a blog post called Breaking the Fourth Wall, about code that talks back. The biggest risk is being spammy or sending people on a wild goose chase, so the feedback has to be accurate and pertinent. Jessica says to test in production as well as before it.
Roni is @DoppleWare on Twitter and writes on Medium, and the favorite board game Roni mentions is New Angeles, though Roni rarely plays the same game twice, since repeat plays become rule hacking.
Jess and Roni talk about what continous feedback: where it came from, what it looks like in the context of a dev proces, and the benefits it can bring to engineers and developers. They also discuss Roni's observability project, Digma.ai... and his other passion, complicated board games.
Matty talks with Dagna Bieda, a software engineer turned career coach with over ten years of coding and more than three years of coaching, about communication, empathy and the people side of engineering. Dagna has coached clients from large companies and small ones, with two to twenty years of experience and backgrounds from self-taught to bootcamp to college. The cold open is Dagna: "It's important to realize that at the end of the day, we're people working with other people to create products for other people."
Dagna says everyone comes to a conversation with assumptions from family and culture, and clarifying them is a key communication skill. There are two kinds: the ones you hold yourself, such as how many people and how much time a feature has, which you can state, and the ones you make about what others assume, where it's better to ask and overcommunicate, because "there are no stupid questions." Dagna recommends starting with an "I sentence" to create a safe space: "I didn't get it. Why don't you give it another go?" so the other person doesn't get defensive. Matty says that when you keep saying people just don't get it, that's on you, and Dagna says you can't control other people's actions but have an impact on how they behave, and if you think you work with idiots, people sense it.
Matty asks how to operationalize empathy, recalling the 2014 devopsdays talks that were all about it and a talk title, "You've Convinced Me We Have to Collaborate Now. How the Hell Do I Deal With People?" Dagna says empathy is misunderstood as having to feel what others feel. What works in the workplace is understanding other people's priorities: an engineer wants to fix bugs, implement features, learn and have fun, while a project manager cares about business value delivered on time with no disruptions and isn't moved by refactoring without a tangible change. Dagna gives clients a worksheet to map stakeholders' priorities, and says pitching an idea by showing how it fits their priorities is "how brilliant engineering ideas get actually implemented." Matty says it's like a talk called The 5 Love Languages of DevOps: a change that makes sense to me might land differently for you, like 30 milliseconds faster matters because of an SLA.
Dagna's own stuck point was assertive communication. Dagna is Polish and direct, and recalls a company-wide meeting of around 300 people after layoffs where Dagna tried to raise concerns with leadership with good intentions, and the boss asked why Dagna had called the entire leadership team idiots, which Dagna says was not what was said but how it came across. The lesson is that intent and how you come across can differ completely. Assertive communication means talking about your needs and wants respectfully so others feel heard. Dagna says probably every senior engineer sounds like an asshole now and then by being too direct, and that humans aren't rational beings who only look at facts.
Matty says systems are complex and have ramifications beyond your microservice. Dagna says engineers are so deep in the woods that they see one tree, while a manager sees the health of the project, a director sees groups of trees and the C-level sees the whole forest, and stepping outside the forest is critical for career advancement: "what got you to that senior position is not going to get you past that position." Once your technical foundation is solid, soft skills matter more. Matty points to earlier episodes on principal engineers and says even engineers who became engineers to avoid people deal with people.
Dagna says the way you think impacts how you act, and compares the subconscious to background programs running on an operating system: people around you growing up "install beliefs" with a sudo command. A common engineer belief is that their work speaks for itself, and it doesn't, since nobody sees it unless they pull the repo. Dagna says marketing your work is a communication skill, and changing a belief is possible because people keep changing. Matty adds that ops roles have the same visibility problem as a corporate lawyer, and recommends knowing how your company makes money. Dagna tells of two features: a refactor that cut a mobile build from about 4 minutes to 20 seconds, interesting but affecting two engineers, and a boring build for a huge client that got praise from the boss's boss, the sales rep and the client. A conversation with the manager showed where the impact was, so ask how features affect the business. Dagna adds that if you can't ask your manager that, run. Matty recalls not knowing for years what the bank's treasury services line of business did until reading its intranet page.
Dagna's first coaching step is figuring out a client's values and the last is finding companies that share them, so you aren't reacting to job postings. Matty notes that interviews are full duplex: the candidate is also evaluating the company, and interviewers are usually untrained. Dagna says candidates who ask questions back are people Dagna wants to hire, and that confidence in your skills lets you negotiate, like the six-month sabbatical Dagna negotiated four months into a startup job, which was spent traveling across Southeast Asia.
Dagna's one recommendation is to ask more questions and put assumptions on the table. Dagna's coaching site is themindfuldev.com, with a case study video, and the two agree to do a follow-up on confidence.
Matty talks with Tim Banks a year after their November 2020 conversation, Breaking Down Gates, recorded around the end of 2021. Tim has changed jobs and now works at the Duckbill Group, and Matty has moved to Pulumi. They revisit what they predicted, and range over the Great Resignation, remote work, who gets hired where, and how to make conferences more inclusive. The cold open is Tim: "Far be it for me to say like I have all the answers. I mostly have questions."
Matty notes that last time they hoped the pandemic was nearly over, and that devopsdays Chicago skipped 2021, the first year without one since it began. Matty hoped organizations wouldn't go back to the old way, and sees some insisting on it, with people not putting up with it. Tim says they'd both predicted that the second people felt comfortable leaving jobs where they were mistreated they would, and they have: people quit and figure it out later, or go to employers who treated people well, and the leverage is with employees now.
Tim says most companies "survived" and didn't thrive, because culture didn't change for remote work. Managers demanded Slack presence and cameras on, and couldn't manage people instead of overseeing them. Burnout and bad management are the leading drivers of quitting, and "retention problems" are a lagging indicator. If you can't staff positions, Tim says, "that is a you problem that you created." Matty adds that culture is a reflection of the behaviors of the people there, and that many leaders have not really tried remote work, since people working from dining tables during a pandemic isn't remote work. Tim says bad management "knows no industry" and may not be malicious, since a great in-person manager can become a poor one when circumstances change. Tim also says companies claiming you can't do culture from a bedroom can change the culture, because technology drives cultural change.
Tim says the pandemic stripped away distractions, leaving "just you and your shitty job and your shitty manager," and that remote hiring let people apply anywhere. Matty asks about traditional enterprises like banks, insurance and manufacturing. Tim calls the Bay Area's mystique dubious and says the pandemic showed talent is everywhere, so influence can decentralize, though money remains concentrated there. Tim says older, risk-averse industries and managers whose careers were built on overseeing people directly are the ones pushing people back, along with real estate commitments, and that pushing people back is how you tell a tech company from a company that uses tech. Matty says most tech workers are at traditional companies, and that people on this podcast forget how small a sliver of the industry works in newer ways. Tim distinguishes tech workers from the tech industry, such as Home Depot's retail versus e-commerce sides, and says the disparity drives people away. Matty calls this bimodal IT, which Matty doesn't think is a good way to run an organization.
Matty describes public-sector leaders who accept that someone might stay only a couple of years, and says a great manager says "you've outgrown this team." Tim, who worked in government contracting, says people in some public-sector jobs are incentivized to stagnate technologically, like protecting a platform you've certified in for decades. Tim says people used to leave for raises, and now they leave to not hate their job, and the only reasons to stay are the money, the job and the manager. Matty shares stories of old systems nobody knows the purpose of and is afraid to turn off, and Tim says people are still patching COBOL. Matty says whatever you're running, you're doing good work.
Tim is glad KubeCon North America 2022 is in Detroit and asks conference organizers to stop picking expensive places like the Bay Area, New York, Napa and Lake Tahoe. Tim wonders why events don't use spaces at historically Black colleges, such as those near the devopsdays Austin venue or in Atlanta or DC, and says virtual conferences let Tim talk to more people than ever, and "I don't want to give that up." Matty agrees it should be both, not just hybrid, since All Day DevOps has always been virtual, and some people prefer virtual for accessibility reasons. Matty says diversity of organizers matters because they may not know about those facilities, and venues are often repeated because "we've done it there before." It's too late for devopsdays Chicago 2022 because a contract is signed, but Matty commits to keep working on it.
Tim says ticket prices could be lower and sponsors could pay for plane tickets, hotel rooms or childcare facilities, and organizers should talk to neurodivergent people, parents and people who can't get days off. Matty says many devopsdays could be run fully sponsor-funded, and that free tickets should not require people to justify why they need one: a checkbox saying "my employer doesn't pay for this" is enough. Tim says having done it, being asked to justify is demeaning. Tim says the industry is more inclusive than it's ever been, which has allowed harder conversations about technology and how workers are treated, but "if you look out in that audience and it doesn't look like something you would consider inclusive, then you have more work to do."
Tim is @ElCheffe on Twitter and at the Duckbill Group.
Breaking Down Gates with Tim Banks (previous episode with Tim)
Matty talks about DevOps and security with Sue Choi, co-founder and CEO of Mondoo, an infrastructure security company, and Dominic, another Mondoo co-founder, co-creator of InSpec and other tools, whom Matty knows from Chef. The conversation is about why security and DevOps still struggle to work together. The cold open is Dominic on attackers: "the hackers are making the same discovery with the same speed," except they're "highly motivated to use them as quickly as they can."
Sue says that if software is eating the world, hackers are having a feast, and there aren't enough security professionals. DevOps teams already do a lot of security work that's hidden. Sue thinks security should be everyone's job but people don't know how, aren't incentivized with shared metrics, and often "black out" at the topic because the stakes are high. Matty cites a tweet, which turned out to be from someone the transcript renders as Cat Sweet, asking why security incidents don't get the same blameless "time I took down production" stories, with the reply that they cost money, to which Matty answers that tech incidents do too.
Dominic says both aim to make infrastructure run as intended: a service that's secure but down isn't useful, and one that's running while giving out credit card numbers like Halloween candy isn't either. Not every security finding will be fixed or needs to be. Sue calls the security team's mandate "paper security," as opposed to real security, which is what DevOps cares about and is hard to determine. Matty says organizations often claim a regulation requires something when it's how they implemented the control, and uses a Simpsons image of layers of physical security ending at a broken screen door.
Dominic recalls that about ten years earlier, at Deutsche Telekom, when cloud and DevOps were new, an auditor arrived with paper security manuals that didn't fit how they ran infrastructure. The team rejected the binders, distrusted security and ended up worse off. It resolved when a new pen tester and auditing team came in with fresh eyes and said some docs applied and others didn't, and together they wrote new requirements and put them into the automation pipeline. Dominic adds that ransomware attacks leverage the same automation that DevOps preaches.
Matty says policies are organizational scar tissue, and that almost nobody blocks you because they're a dick. Sue says security conferences lack empathy and team-building workshops, but that's shifting as security diversifies. Matty says classic security is adversarial, which shapes the whole outlook, like Matty's own old-school ops view of developers as the enemy. Sue says CISOs often feel peers dislike them, since CTOs can relate work to business value while security talks about risk, so the change needs to start at the top. Matty says ops is like being the corporate lawyer, known only for failures, and Sue wants to invite security to devopsdays.
Dominic says both sides have tried crossing over, but each lacks the other's context. Security asks why ops can't just auto-remediate, and ops worries about what that does to infrastructure. DevOps has built the muscle for safe, reproducible change with development teams, and it's time to extend it to security: don't hand over a list of 300 findings, give 10 critical ones to tackle together. Sue says that needs negotiation skills and a decision-making framework. Matty says shift left without nuance sounds like developers doing all the security, like NoOps, which failed because domain expertise matters, and Dominic adds it's subject matter expertise around one table.
Dominic suggests picking one topic, like SSL/TLS, and discussing it with the security team, and not taking all 300 results. Sue says people in security have a hard time, like a SOC 2 checkbox about cameras at the entrance for a fully remote company, and DevOps people's answer is always "it depends." Dominic describes two trust-breaking stories: security freezing an environment for three weeks for an audit, and the DevOps team changing the environment the Monday after certification so nobody knew what happened. You can't change the other side but can change yourself.
Sue would love success stories to match DevOps metrics. Matty agrees and says that talk-worthy failures can be told without revealing attack details. Sue raises Equifax, and Dominic says there's technical security and legal security, and to find the security person who will have the technical conversation.
Dominic describes the supply chain as everything that goes into building software: dependencies and infrastructure components. Shift left applied to supply chain with the same 300 findings means people react only to critical issues, and then findings get relabeled as critical. Tools are full of false positives, and "nobody's going to give me those 15 minutes of my life back." Dominic says progress comes in three steps: visibility, prioritization, then fixing, and thinks auto-remediation vendors are getting ahead of themselves. Matty notes that requiring CIO approval for every open source component isn't protection.
Dominic's two suggestions: build a relationship with a security person, even by talking about anime or Star Wars, and get visibility into your infrastructure, since companies have a bigger visibility problem than they admit. Sue's are to make time to build relationships in a remote world, and to align on goals, since one DevOps team assumed security should own policies and tools while security didn't know.
Matty talks with nine organizers and speakers of the second Deserted Island DevOps, the DevOps conference held in Animal Crossing, which Matty calls a personal favorite virtual event. The organizers are Austin Parker, a developer advocate at LightStep, and Katy, a technical community manager at CircleCI who emceed. The speakers are Ana Margarita Medina, a senior chaos engineer at Gremlin, Rin Oliver, a technical community builder at Camunda, Angelina Uno-Antonison, a software architect at the University of Alabama at Birmingham School of Medicine, Arri Blais, a senior software engineer at Embark Vet, Jack Knives, lead DevOps engineer at Moda Operandi, Laura Santamaria, a developer advocate at LogDNA, and Serena Tiede, an SRE at UnitedHealth Group. The cold open is Katy: "I just love hyping people up, Matty."
Austin says the idea was the same: a virtual event in a deliberately constrained virtual space, with the constraints meant to inspire creativity. Austin announced it way ahead of time so fear of failure couldn't talk them out of it. This year added more production value, artwork and merch, and turned last year's charity shoutouts into a fundraiser for The Trevor Project, with a ticker during the event. Austin says it raised over $6,000, though the turnout wasn't the hit Austin expected. Angelina says the ticker made the event mean something more, and Austin says the commitments of no entrance fee, free on Twitch and accessible with transcription are founding concepts, and that getting all those people together means you can probably ask them to give.
Katy says as emcee there's no off mode, and introducing people talking about what they're passionate about brings joy. Katy ran a dedications line in a Discord channel, reading people's shout-outs aloud in a radio voice, since Katy once wanted to be a radio DJ. Austin also says introductions use fun facts in place of bios, which was accidental last year and is now deliberate, since a first-time speaker's short bio next to someone with a page of accomplishments invites implicit bias. Matty says devopsdays Chicago stopped having MCs do intros for a similar reason, since MCs know some speakers and not others.
Ana, a first-time speaker, only started Animal Crossing around November or December of 2020, and had trouble with the abstract. Ana zoomed out from reliability to interpersonal things, like checking on neighbors and building trust and safety in teams, and tied it to reliability with the example of buying two axes because your tool will break. Rin's talk was on community and resilience, arguing resilience means having the tools to handle pain in a way that suits you and has to be worked at, like Roald, a villager who loves working out. Rin liked that the event was on Twitch in Animal Crossing and not just Zoom.
Angelina spoke about Apache Kafka, themed with Animal Crossing and a dinner party, partly to make it less intimidating for the computational biologists at a monthly meeting, and partly because burnout had made Angelina not want to give talks, so a video game theme was a way to do it differently. Based on feedback, Angelina will keep using it. Arri gave a talk on getting started with accessibility advocacy, since people often don't know how to begin, and liked that the conference mixes soft skills and technical talks. Matty and Arri compare single-track and multi-track conferences. Jack, a first-time public speaker who submitted a throwaway proposal in two minutes, spent two weeks on drafts and dry runs and an hour in GIMP making slide patterns, and is proudest of the line "Blathers was the developer because he's afraid of bugs." Laura's talk covered the dangers of third-party tooling, and Laura live-tweeted the whole event. Serena planned a technical talk about running Jaeger that turned into one about making friends with tracing.
Austin says the production lesson is to tell speakers which side of the screen to stand on: the scenes assumed people would stand on the left, the first speaker stood on the right, and everyone mirrored, so Austin was furiously panning the camera. This year the slide inset was wider, about 800 pixels against 600 the year before, and speakers were asked to make text bigger. Matty suggests a set of best practices for speakers, including having someone else run the reactions. Laura says the Q&A after a prerecorded talk, with the speaker still in Zoom and Katy relaying questions from Discord, felt more like in person than typing into a chat. Laura and others note that participants were engaged, with watch parties on other islands and new channels, including one that spun off an ADHD ops discussion, which Serena says a platform like Discord handles better than others.
Katy says last year had energy from the start of the pandemic, while this year the talks were more needed, messages of wholesome resilience. Austin says it's a story about resilience and shares something not widely known: LightStep had layoffs on April 30, 2020, the day of the first event, and Katy was affected and was offered an opt-out of hosting. Katy says, "don't you dare take this from me," and that Deserted Island got Katy through that, gave Katy a new community and friends, and "pretty sure it got me my next job." Austin says it was one of the hardest days professionally, and that this year Austin repeatedly doubted doing it again because the mood had changed, but decided to "take the limitations, take the things that suck and compress them into a tiny ball of suck, and then try to put that tiny ball over to the side and fill that space with good."
Serena and Arri's advice for other events is Discord's hallway-track feel and building a helpful, wholesome community. Katy asks people to watch the YouTube videos and follow the speakers on Twitter.
Melody: melody.dev
From the publisher's feed