Arrested DevOps

Arrested DevOps

By Matt Stratton, Trevor Hess, Jessica Kerr, and Bridget KromhoutTechnologyTech News
Download on the App Store

Arrested DevOps episodes

  • The Meltwater Transformation

    In this episode, host Jessica Kerr is joined by guests Gene Connolly and Joan Freed to discuss DevOpsiCon and their DevOps transformation at Meltwater.

    DevOps at Meltwater

    The panel discusses how DevOps has transformed the working environment at Meltwater. When Joan started at Meltwater, engineering and operations were kept entirely separate.

    Joan: “There was this huge wall because engineering knew how things were built and operations knew the production systems...but there wasn’t any cross-pollination there.”

    The lack of communication between the teams caused friction and slowed down product delivery. The process could take as long as six months.

    Joan: “We’re a Software as a Service company so we always want to build products that are sticky and keep our customers around.”

    Joan discusses how Meltwater began to move to a more agile business model. Getting engineers and operations people in the same physical location to work this out was important. Creating truly cross-functional teams couldn’t happen over Skype alone. DevOpsiCon was born to facilitate this transition.

    DevOpsiCon!

    Buy-in wasn’t immediate throughout the firm. Teams were updating legacy systems throughout the company and some were more focused on work than enablement.

    Joan: “The first one was a little bit sneaky. The next one was in Manchester and that one was a little less sneaky.”

    The panel discusses the importance of investment in enablement. This allows the feature teams to focus on the features they were actually building instead of all the underlying infrastructure.

    Joan: “Having them build everything they need is not realistic.”

    The fact that Meltwater has offices all over the globe makes any wholesale shift of company culture difficult. The panel discusses the need for face-to-face interaction to facilitate change.

    Gene: “There are these events like DevOpsiCon where we solve the problem by bringing a broad set of people together.”

    DevOpsiCon has grown beyond just developers. Product, UX, support, and even sales have joined the event to help build relationships and give insight.

    Jessica: “That is very DevOps in the sense that DevOps says that you need to take these concerns that cannot be separated...and you can’t put those responsibilities on different teams and make them fight.”

    Unconference

    The panel discusses the value of DevOpsiCon as an unconference. Attendees collectively set the agenda on each day of the event.

    Gene: “That format has been so much fun and created so much value.”

    Gene: “The event becomes the conference you didn’t know you needed at the start. It adapts collectively to what the organization needs at that moment.”

    Ongoing Transformation

    Gene discusses how allowing teams the freedom to experiment can yield results that a top-down structure could not provide.

    The panel talks about the structure of mission, framework and support teams.

    Joan: “By keeping things loosely coupled, it gives teams a lot more freedom to make the technology choices that are best for them.”

    Ownership Challenges

    The panel discusses the ownership challenges of moving from the datacenter to the cloud specifically and large enterprise changes more generally. The investment in enablement is critical.

    Joan: “There’s very much a not built here mentality that can take place.”

    Gene: “Software can benefit from being owned by many people.”

    The panel talks about overhead costs to both software and human changes.

    Gene talks about the challenges of moving to a full mission responsibility model.

    Gene: “You need to find efficiencies more effectively than you’ve ever had to find efficiencies before.”

    So! Has It Worked?

    Gene: “Oh wait, you’re looking to me for an answer?”

    The panel discusses the results of the ongoing DevOps transformation at Meltwater. The six month process for production changes has transitioned to a constant stream. Beyond the productivity, it has also been a quality of life improvement for the employees.

    Gene: “We’ve got a Slack channel for all the production changes and it is constant activity in it all week long of changes going out.”

    Joan: “We’re now getting hundreds of features out every six months to our customers.”

    The panel discusses the ways in which transformation is, or should be, a constant process. The change goes on.

    Find out more about Meltwater, their DevOps transformation, and the process of an unconference on their blog.

    Episode art by Evelyn Kerr.

    1 hr 1 min
  • DeliveryConf

    Matty talks with two organizers of a brand new conference, DeliveryConf, which runs January 21st and 22nd, 2020 in Seattle. Ken Mugrage is a Technology Advocate at ThoughtWorks who works on DevOps and continuous delivery and helps organize devopsdays Seattle. Sasha Rosenbaum is a Program Manager on the Azure DevOps team at Microsoft and co-organizes devopsdays Chicago with Matty, who says this isn't paid programming. The date was chosen to avoid the heavy conference season in a city where they probably won't have snow.

    The Gap Between Advice and Implementation

    Sasha says devopsdays keeps having non-technical conversations, partly because the audience is so mixed that you don't know who is dev, ops or PM, and partly because as an open source event it doesn't want to endorse products or go deep on technology. People leave with amazing ideas and "have not seen anybody do any implementation of stuff." Ken says every conference has a DevOps track, but the talks say to do security testing, not how it was done on a specific project. Looking around, there were DevOps conferences and language conferences "but not CD conferences."

    Ken picks on the talks Ken gives, which stay at a high level because the audience's tech stacks are unknown and 30 minutes isn't enough to go deep. The hope is that what was 5 minutes of a talk becomes a 30-minute talk on how compliance was done on one specific project. Sasha adds that the conference aims to be vendor-neutral, so people can showcase products and techniques and compare them. It's for hands-on practitioners, closer to "hone your craft" than an introduction, and Sasha says a first-timer is probably better off at devopsdays. One gap Sasha hopes to close is continuous delivery for databases: "you can't really continuously deliver your entire software application if you're not able to continuously deliver your database."

    Discussions as First-Class Content

    The format is a 30-minute talk followed by a 20-minute facilitated discussion, which will be recorded. Sasha says they love open spaces and wanted to make the conversation first-class content, so people who couldn't attend can listen. Ken says the conversation is about the topic, not the talk, and the aim is "to take the speakers off the pedestal," though the speaker can join or not. The facilitator will ask what challenges people have, what successes they can share, and the magic wand question: what would you want to see in the next two to five years, in product or culture, to make this easier. Sasha notes that Q&A isn't for arguing with the speaker, and this gives people with a different opinion a place to voice it, with rooms to continue offline.

    The recordings are audio only, to respect privacy, with a microphone used as a totem, and participation is optional. Ken's hard metric for success is that the discussions get at least as much traffic as the talk videos. Sasha's softer one is seeing the light bulbs go on and people learning they're not alone in their experience. Ken adds that it's a "paid focus group after every talk" for vendors.

    Program

    The tagline is learning from today and shaping tomorrow. The opening keynote is Jez Humble and Dave Farley, co-authors of the book Continuous Delivery, which has its 10-year anniversary next year. Heidi Waterhouse from LaunchDarkly will speak on feature toggles. Ken teases machine learning in pipelines and hands-on security sessions, and a second-day panel of executive-level people on what CD will be in the future. The CFP closes the night of the recording.

    Fast Versus Safe

    Matty asks where the state of the art is. Ken says the industry is at a crossroads: there's a focus on how long a pipeline takes from commit to production, and a 3-minute pipeline can't have done compliance and security checks. Jez's definition of continuous delivery on continuousdelivery.com "specifically says safely," and Ken was glad DORA now says "on-demand" for the top tier. Matty's version is that it's not as fast as possible but as fast as you need it to be. Matty also tells of an e-commerce boss who said the target for a site performance dashboard was "not slower than last month," which is good while you're improving and less useful once you reach what's appropriate.

    Ken argues that what you measure isn't the business value of deploying on demand but the value of the thing you deployed, framed as a hypothesis, and wonders whether machine learning in the pipeline could learn what a change did to the business and kill the canary automatically. Ken also says "I don't think hardly anything that we call AI is AI." Sasha says not everyone is a startup deploying straight to production: people have different compliance requirements and levels of legacy, and "I don't even like the word legacy because like, hey, this is software that's making money for you." Matty supplies "heritage systems."

    Matty brings up incident response, where you'd like to restore service without bypassing your normal checks. Ken describes a pipeline on the most recent team that was fairly long with all the security checks, but had a short circuit for emergencies. Basic unit tests and fast tests ran and produced the installer, and someone with the right permissions could click a button to reach production, with the rest of the pipeline tailing it. It was still the same pipeline with things skipped.

    Tickets, Sponsors and Who's Missing

    Tickets are at deliveryconf.com, and the code ADO gets 10% off. Ken says the price, about $400 to $450, is higher than most community events because of three tracks and all the recordings, and people who can't attend for financial reasons should reach out, since some sponsors may help. The conference is a not-for-profit backed by devopsdays Seattle, and Ken says it lacks the budget for big promotions. Sasha says hearing that people who didn't know of Sasha's involvement call it a solid conference means "we're actually hitting the mark."

    Gold sponsorship includes a sponsored talk, and the goal is deep technical talks and not product pitches. Sponsors submit decks two weeks ahead for review. Ken tells sponsors to bring engineers and product managers, not just sales staff, and says the aim is real solutions to real problems, without requiring live coding. Sasha closes by citing a study that only 3% of women identify as technical in the Dev/CICD space, says it's proving hard to find non-male speakers for the CFP, and invites diverse participants onto the stage.

    Matty Stratton talks with guests Ken Mugrage and Sasha Rosenbaum about their new event DeliveryConf.

    35 min
  • devopsdays Cape Town 2019

    Bridget records on day 2 of devopsdays Cape Town 2019 with two organizers and two speakers. Cobus Bernard is a technical evangelist at AWS, and Adrian, who does DevOps-type work at Salesloft, co-organizes with Cobus. This is the fourth year of the event, which came together after Bridget visited in March 2016 and held its first event that November. Daniel, who works at Datadog, was the opening keynote speaker at that first event and is back to speak this year. Devi Moodley heads up the DevOps team at Nedbank, one of South Africa's four biggest banks.

    Transforming a Bank Founded in 1888

    Devi's talk went straight to the practicalities of transformation at a giant organization. Devi says "we were not a Spotify": Nedbank is a traditional bank established in 1888, with traditional values and local and international regulation, which makes bringing in "a sense of looseness" very difficult. The approach was to embrace those traditional concepts, accept who they were, and start from there. A bank has to demonstrate value before anyone adopts anything, so they went in with evangelists who had credibility and used them as poster boys for the message. Devi says it's about generating financial value, doing things cheaper and better, and that as long as they generated value they could get more investment. That is how they're growing "slowly but surely, one pipeline at a time."

    Adrian says it's easy to find good stories from unicorn companies, so the program needed balance, and if a big enterprise known for moving slowly can adopt DevOps, "it's possible anywhere." Devi's bank is a sponsor and has spoken at the event for the past few years, and Adrian notes many South African banks have adopted DevOps. Bridget says Minneapolis likewise has people from large local retailers, since it isn't only for the tech vendors.

    What MMA Taught Me About Working in Tech

    Daniel's talk title was What MMA Taught Me About Working in Tech. Computers and mixed martial arts have both been big parts of Daniel's life, and Daniel realized that a lot of how they act, make decisions and interact with people came from martial arts training. Cobus says balancing technical and cultural talks is always fun, because "tech is never the problem. It's always people," and the organizers were surprised how much culture showed up even in talks they expected to be technical. Bridget says that's also true of open source: adoption depends on the people deciding whether to trust a project, and a technically excellent tool with no documentation and six months of silent pull requests looks risky.

    Adrian agrees that if a project's pull requests sit for months and the last commit was six months ago, "if there's no community around it, it's not going to happen." Daniel pushes back that vitality is worth checking but doesn't settle it, since some software Daniel uses hasn't been updated for 10 years because it still does what it was designed to do, unless there's a security problem. Daniel adds that tooling is getting better at flagging that, and suggests building checks into pipelines or occasionally looking at the CVE list.

    How a Bank Picks Its Tools

    Bridget asks how Nedbank decides whether an open source tool is okay. Devi says open source is allowed in principle, subject to tight security checks. The method starts with the community that will use the tool: unpack their problems, map the whole value stream, have the team prioritize the biggest issues, then look for tools that would solve them. A proof of concept runs for open source and paid tools alike, and the outcomes are evaluated against the original problems. Where timing matters they "literally stopwatch our developers" before and after a learning cycle, rate the tools on a Harvey Ball scale of 1 to 4, and sometimes test two tools in parallel with different teams.

    An easy tool takes about two weeks, starting at the beginning of a team sprint after a few days of training, and a heavier one about six weeks. Sign-off takes longer than the proof of concept "because we're a bank." Devi says a full process map can be done by locking people in a room for four hours. Bridget likes that it happens inside the sprint, since for the period of the decision it has to be the work, and Devi says that makes it practical and not "a science experiment scenario." Daniel says the depth of the answer was impressive and unexpected.

    Cape Town, Joburg and Who Runs the Conference

    Bridget asks about the link with the Johannesburg community. Adrian's involvement has been mostly Cape Town, and jokes that Capetonians give Johannesburg a hard time because "They don't have the ocean." Adrian says a growing DevOps community exists in Johannesburg and the Cape Town team should help them start a devopsdays without running it. Adrian mentions the Python community, where PyConZA needs meetups in both cities to qualify under the Python Software Foundation.

    Daniel says devopsdays events are meant to be local, and that in 2013 when the first Paris event ran, people traveled from all over Europe because it was the only game in town, but that was never the goal. Daniel says roughly 75 events were planned for 2019 with no central committee authorizing them, and that if a team in Marseille wanted to run one, Daniel would share the mistakes made and then back away. Bridget relays how Matty Stratton used Minneapolis, which ran a few weeks ahead of Chicago, as a model to copy what worked and skip what didn't. Adrian adds that cross-pollination across organizations means Johannesburg could eventually help Cape Town improve too.

    Running a Conference Out of GitLab

    Cobus asks Adrian to walk through the process used to run the conference. Adrian says nothing is public, and the idea was to use DevOps to run a DevOps conference, after taking a management job and starting to care about project management. There are private GitLab repositories, one for the meetup and one for the conference, each with a wiki as a knowledge base (venues, sponsors, potential speakers) and issues for actionable items. The meetup gets one issue per month from a template checklist, covering the venue, time slot, meetup.com page, sponsors and announcements on Twitter.

    For the conference they use a milestone per year and an issue per sponsor, labeled as a lead until they sign, with templates for each sponsor level because, for instance, a platinum sponsor gets to choose the Wi-Fi password. The same idea covers speakers: whether their bio picture, abstract and Twitter account are in place. Adrian says it's still a work in progress, and the stated goal is to automate it so a script can run next year's event. Bridget plans to deputize Adrian to help rewrite parts of the organizer guide.

    Takeaways

    Daniel's is a talk by Rory, who works at Microsoft, on building things usable by people regardless of their abilities, an area Daniel says they don't know much about but have some personal experience with. Adrian heard the talk earlier in the year at DevConf and says it made the point that building accessible websites is easy and "it's no longer an excuse to have a website that's inaccessible." Devi's takeaway is from Daniel's talk: "the success of any organization is based on the well-being of your people," and work-life balance "isn't just a couple of words on a piece of paper."

    Adrian has a list of 10 things devopsdays could do better, and most aren't about Cape Town, so Adrian may aim at improving the global DevOps community. Bridget, who says they're the current titular lead of the global devopsdays community, would like to subscribe to that newsletter. Cobus noticed new people asking questions in the open spaces and more experienced people helping, and says the community should invest in making it easier to ask, since many people take remote jobs and disappear from the community scene. Bridget closes by praising the way the organizers pre-selected open space ideas, gave participants a web interface to add to them and sorted them ahead of time, a format Bridget hadn't seen before that got discussions going right away.

    Bridget chats with Devi Moodley, Daniel Maher, Adrian Moisey, and Cobus Bernard at devopsdays Cape Town 2019.

    • Devi Moodley's talk: From Science Experiment to Enterprise Rollout
    • Daniel Maher's talk: What MMA taught me about working in tech
    • Community & Event Stuff

      If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf

      Upcoming conferences
      Open CFPs
      • lots of DevOpsDays
      • Discount codes
        • ADO2019 for 20% off lots of devopsdays
        • Velocity Berlin Nov 4-7 2019 - discount code "ADO2019" gives 20% off for Gold, Silver, and Bronze passes.
        • 37 min
        • State of Devops 2019

          In this episode, hosts Jessica Kerr and Matty Stratton are joined by guests Dr. Nicole Forsgren and Mr. Jez Humble, two of the authors of the 2019 Accelerate State of DevOps report.

          Historical Context

          The panel discusses the history and purpose of the report. This is the sixth year it has been produced. How did the report start and what questions is it seeking to answer?

          Nicole: “Once upon a time, kids, making software was sad. Making software used to light people on fire!”

          Nicole: “So why do we even do this DevOps thing? It’s because we want to make that software process easier and better.”

          Jez: “These approaches have worked even in highly regulated environments. They’ve worked everywhere.”

          Nicole explains that capabilities and practices are more important than tools.

          Nicole: “There’s no such thing as DevOps in a box.”

          The group discusses the limitations of systems data vs survey data, and the importance of collecting both. Survey data is low resolution but high signal and vice versa.

          Jez: “Over-precision is something that’s a real problem in our industry.”

          Nicole: [“How to Measure Anything” - Douglas W. Hubbard] ( https://www.amazon.com/How-Measure-Anything-Intangibles-Business-ebook/dp/B00INUYS2U )

          The panel discusses the importance of starting from hypotheses instead of looking for any spurious correlations that exist in highly related systems.

          What’s New This Year?

          The panel discusses some of this year’s findings. The “retail apocalypse” over the last decade has had some surprising effects. Regulation has less of an impact than many industries assume.

          Jez: “We do see high performers in large companies who are highly regulated.”

          Matt: “People will sit there and say well, I have over 5000 employees and I can’t DevOps so now I have an excuse, and that’s not the case.”

          Enterprise!

          Change management processes are one of the key areas of difference between large and small enterprises. Jez discusses how to make that process more lightweight even in a large organization.

          Jez: “Risk is also about upside risk. If you can’t move fast at delivering software, that’s a risk to your business.”

          The panel discusses the importance of keeping a holistic view even in a large enterprise with specialized roles.

          Nicole: “You go from change management theater to strategic change management for a massive organization. It’s scary but it is also dope!”

          The report has found year after year that a slower change management process can paradoxically result in more instability, not less. The panel discusses the reasons this experiment was not successful and how enterprises can implement more nimble processes going forward.

          Productivity

          The panel talks about what productivity actually means. The report uses an interesting definition: “Productivity is the ability to get complex, time-consuming tasks completed with minimal distractions and interruptions.”

          Nicole: “You may be just closing the tickets that are meaningless but are easy to close.”

          Research shows that real productivity reduces burnout and improves work-life balance. The group discusses how to separate that real productivity from gamification and busywork.

          Jez talks about scaling up and how centers of excellence may not be as useful as previously thought. Nicole points out that the report has given a lot of starting points for people to begin making changes in their organization.

          Remember to read the report!

          53 min
        • devopsdays Chicago 2019

          Matty and Trevor record live at devopsdays Chicago 2019, the sixth time the event has run, with three speakers from the program. Jessie Frazelle, who is self-employed, talked about open source firmware and roots of trust and why they matter. Veronica Hanus, a contractor in Brooklyn, gave an Ignite talk on how developers' attitudes toward code comments affect growing developers and, later, their documentation styles. Jeff Smith, Director of Production Operations at Centro, talked about why ethics in technology is difficult. It is Veronica's first devopsdays ever and Jessie's first Chicago. Jeff says it's the fifth time at devopsdays Chicago, and that an open space on ethics the year before helped inform the talk: "a conversation that needs to continue to happen and grow."

          First Impressions and Fifth Ones

          Veronica arrived wondering how public transit works outside New York, and found a team that helped with transportation, served "the best food you've ever eaten at a conference," and even asked speakers what song they wanted when walking on stage. Jeff says the event has grown and the audio and video have gotten smoother, and calls it "the bar at which I judge all of the other conferences." Jessie praises how well organized it is and the varied audience, including a data scientist and a technical writer, as a way to avoid groupthink. Veronica's favorite conversations were in open spaces about bringing the lived human experience into a talk submission alongside the technical content.

          What Could Go Wrong

          Matty sees a parallel between Jessie's talk and Jeff's, since it may have seemed a good idea to put a web server somewhere it didn't belong, and asks how these things can be misused. Jeff says we misapply technology ourselves in operations and engineering, choosing tools because they're interesting without thinking through downstream consequences. The example is the recent Zoom breach, where a local web server was meant to make reinstalling the software easier. Jeff recalls Google's image-classification mistake involving Black people and says "there wasn't a Black guy in that room," so the fix is broader perspectives and thinking one step beyond the task in front of you. Trevor proposes asking "what could go wrong" seriously instead of facetiously, and Jeff says an IAM policy statement ID Jeff saw in AWS was literally "what could go wrong," with allow star star.

          Jessie says many employees are process-oriented, so if asking the question isn't in the process, you do A, B, C, D and finish the task. Veronica says it comes down to psychological safety, and the innovation possible where it's okay to ask why something is the way it is differs greatly between organizations. Trevor ties it to Jeff's morning idea of an ethical body for technology: people feel pressure to finish but rarely the moral pressure to challenge getting it done. Jeff adds that standards and culture move the bar, recalling a job 10 years earlier where a password in a config file wasn't crazy, until a later job where the reaction was "what's wrong with you?" Jessie says it's culture: if the culture is to make money even at the expense of exploiting customers and nobody faces ramifications, "that's the culture that you set."

          Veronica remembers a government research center that wouldn't put anything in the cloud, where losing data was an expected part of the process: "You lose data, you grieve a little, and then you put your head down and create it again." Looking back, Veronica says, it sounds bizarre.

          Building Psychological Safety as a Leader

          Matty asks Jeff, as someone who leads teams, what to do. Jeff's first answer is vulnerability and being unafraid to say "I don't know." On Jeff's team that's an okay answer, just not the end of the conversation. Jeff also pushes education as part of the job, asking why people who are already on call should learn a new technology on their own time, and says spending an hour or two a week reading in the kitchen is fine. The third is creating space to dissent constructively, so people who disagree know "we're still a team."

          Veronica says the difference between helpful and unhelpful workplaces was time to learn and how things went wrong. In a research job, a machine had a problem nobody noticed for a month, and the team traced it back in what Veronica calls a blameless way, and everyone left knowing how the machine worked inside and out. Veronica says both fields produced the same pattern: learn at work instead of being burned out in the evenings, and when something goes wrong, focus on the process more than the person who made the commit.

          Influence Without the Title

          For an individual contributor without that safety, Jessie says candor with leadership can go one of two ways, depending on the person. Jeff says to think of leadership as a role and not a position, and that whoever is the emotional leader of the team can provide some psychological safety cover, such as going to the manager to say something isn't working, or inviting a quiet teammate who is stewing over a technology to speak. Veronica adds "praise in public, reprimand in private," and to remember a manager may feel just as vulnerable, so pick a moment instead of ambushing someone in the hallway.

          Matty brings up a talk by Anwen Simmons a couple of years earlier on lending privilege, and says one mark of a team with a high level of psychological safety is that everyone speaks about equally. Jeff says that as experience and confidence grew, calling out a bad idea got easier, and that if fired tomorrow Jeff is confident of getting another job, so using that to shield someone else can be huge. Jessie agrees, having started out unable to be aggressive and get away with it. Trevor adds that lending privilege doesn't have to happen in front of the person: talk about their idea in conversation and remind people whose idea the team is now following.

          Key Messages and Plugs

          Jeff's one message is that we should take responsibility for what we create and own the consequences, because the stakes get higher as technology advances. Veronica's is "Comments are a form of documentation, and documentation is good." Jessie's is that people working on different layers of software should talk to each other more and actively listen, to solve problems between the interfaces.

          Jeff is writing a book tentatively called Real World DevOps, written from the perspective of an individual contributor in imperfect organizations, due out in 2020 sometime. Veronica announces, for the first time publicly, the Quick Developer Guides YouTube channel with two other first-time speakers, a series of 10-minute videos on topics from getting started with version control to budgeting for a career change into development. Jessie, who is not actively looking for a job, plugs two books, Soul of a New Machine and Ben Horowitz's The Hard Thing About Hard Things. Jeff says, "She's not looking. She's choosing."

          Matty and Trevor chat with guests Jessie Frazelle, Veronica Hanus, and Jeff Smith in front of a live studio audience at devopsdays Chicago 2019.

          Community & Event Stuff

          If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf

          Upcoming conferences
          Open CFPs
          • lots of DevOpsDays
          • Discount codes
            • ADO2019 for 20% off lots of devopsdays
            • Velocity Berlin Nov 4-7 2019 - discount code "ADO2019" gives 20% off for Gold, Silver, and Bronze passes; early price price ends 20 September.
            • 31 min
            • Protocols and Sympathy with Martin Thompson

              Jessica Kerr talks with Martin Thompson, known for the Disruptor and for mechanical sympathy, about performance, queuing theory and protocols, and how the same mathematics and etiquette apply to computers and to people. Martin spends days on distributed and concurrent systems, helping clients waste fewer CPU cycles, which often means helping people, because "getting people to behave well usually gets software to behave well." The cold open is Martin: "Computers, you don't have to be nice to them, but actually being nice to them, you get better things out of them."

              Being Nice to Systems

              Martin's way of getting people to behave is to understand their motivations, give them interesting things to do and show them value, the best way being to let them meet real users, since so much software goes over a wall. Martin is hired for performance and finds it leads to delivery: you can't know you made something faster without tests and measurement, which needs CI and continuous delivery, so short cycles matter. Queuing theory and Little's Law are the fundamentals behind lean delivery, and "System being code or system being people, it's still the same thing. Same mathematics apply." Mechanical sympathy, Martin notes, is a term taken from racing driver Jackie Stewart, and is really empathy.

              Slack, Utilization and the J Curve

              Martin describes the scientific method as the learning cycle: have an idea, design an experiment, analyze the results, and go back if wrong. The faster the cycle, the quicker you learn. Buffers exist because components run at different paces, and a team with no slack has no buffer to react, so backlog grows. Martin's example is a service that takes 100 milliseconds per job with jobs arriving at one per second, at 10% utilization. At five per second it's 50%, and as utilization climbs, response time follows a J curve, so past about 70% queues form. At 9 jobs per second, 90% utilization, a job waits about a second, and cutting the job to 50 milliseconds makes the system 20 times more responsive, not twice. That holds with the same resources and arrival rate, because it reduces utilization.

              Adding resources only works if the work can be distributed without contention. Martin cites the Universal Scalability Law, which includes the coherence cost, the time to reach agreement, and says Brooks's Law in The Mythical Man-Month is the same math: adding people to a project costs time bringing them up to speed. Martin quotes Rob Pike: parallelism is doing multiple things at the same time, and concurrency is dealing with multiple things at the same time, which requires coordination. A team needs steady state, since Little's Law assumes it, and reshuffling teams constantly changes all the parameters. Conway's Law means dysfunctional team communication produces dysfunctional software.

              What a Protocol Is

              For Martin, a protocol is "just the rules for engaging or interaction between components in a system," whether people or software: the etiquette and the precedence, meaning behavior in order. In real life we deal with out-of-order events by handling exceptions and not expecting perfection. Martin says English is directive with little redundancy, while languages elsewhere approach things from several angles, which is less efficient but safer. Claude Shannon's information theory says communication isn't complete until feedback confirms reception, yet software is often built as if delivery just happens. Martin calls two-phase commit a protocol sold as a silver bullet that is fundamentally broken.

              Zombies, Versioning and Idempotency

              Martin says "nodes just dying is not a problem. Zombies are the real problem," and "Partial failure is much, much worse," so shoot a node in an indeterminate state and move on. Versioning helps at all levels: messages, protocols and state, since a process waking from a long GC pause may act as if no time has passed, and old messages from an earlier session can turn up in a new one, which Martin ties to TLS 1.3 fast-restart replay attacks. At the application level, give every message a unique, monotonic sequence number or correlation ID so duplicates, replays and out-of-order messages can be ignored, as in a banking system with transaction IDs. Martin calls this hygiene, like "a surgeon will not consider performing an operation without washing their hands," crediting Florence Nightingale, a statistician who popularized the pie chart to show infection rates. Tests, CI and monotonic sequences save time in the end.

              Development as a Protocol

              Martin treats testing order as a protocol of precedence: write the test after fixing a bug and it may be bogus, while writing it first and watching it fail gives falsifiability, which is scientific maturity. Jessica adds "Never trust a test you haven't seen fail." Martin says software has only been around a few decades, without generations of trade knowledge, so we should admit mistakes and shorten feedback cycles, and that every protocol needs a way to change, which is why protocols need versioning. Legal systems are codified protocols that change as we learn.

              Amplification, Buffering and Canaries

              Martin says, citing Dijkstra, that software is so novel that metaphors break down, and that one small change, such as a single incorrect bit, can have a catastrophic effect, more than almost anything in human history. Jessica notes that software amplification can be nearly instantaneous, while human systems take time to propagate, and Martin says buffering contains change, comparing it to shockwaves through air versus liquid. Martin gives a historical example of a society that ran by committee in peace, which slows change on purpose, and appointed a war leader only in wartime. Jessica mentions a Cloudflare outage from a pathological regular expression pushed globally, and Martin says use isolation and suitable buffering, like a canary, and Jessica says the same works for trying a process change on one team.

              Slowing Down

              Martin's favorite thing learned that year was an article saying people are always trying to do the right thing, including when procrastinating, which usually means insufficient information or a worry ahead, so it's a canary: "don't treat anything as bad behavior. It's just interesting information." Martin's advice: "just slowing down and pausing is actually one of the best ways to speed up."

              • Martin’s talk on protocols from J on the Beach
              • Mechanical sympathy:
                • https://mechanical-sympathy.blogspot.com/2011/07/why-mechanical-sympathy.html
                • https://groups.google.com/forum/#!forum/mechanical-sympathy
                • https://dzone.com/articles/mechanical-sympathy
                • Universal Scalability Law, Queuing theory, Little’s Law
                  • http://www.perfdynamics.com/Manifesto/USLscalability.html
                  • https://blog.acolyer.org/2015/04/29/applying-the-universal-scalability-law-to-organisations/
                  • https://www.infoq.com/presentations/little-usl-scalability-performance/
                  • http://perfdynamics.blogspot.com/2014/07/a-little-triplet.html
                  • http://www.vissinc.com/2012/09/07/littles-law-isnt-it-a-linear-relationship/
                  • https://medium.com/@__bbak/dont-be-fooled-by-littles-law-18e18dba3717
                  • Florence Nightingale and Pie Charts
                  • Dijkstra on the radical novelty of software
                  • Image credit: Ylva, Ebba, and Kashti Grimm
                  • Community
                    • Velocity Berlin Nov 4-7 2019 - discount code "ADO2019" gives 20% off for Gold, Silver, and Bronze passes, and Best Price ends August 2.

                    • For any devopsdays, try the discount code ADO2019!

                    • 50 min
                    • Pushing Left with Tanya Janca

                      Matty talks security with Tanya Janca, a cloud advocate at Microsoft who went from software developer to security person to cloud advocate, doing web app hacking and incident response along the way. The episode covers threat modeling, the Pushing Left blog series, serverless security, what developers and ops people should know, and the Mentoring Monday hashtag. The cold open is Tanya's invitation to developers: "Come on over, bring coffee, we will worship you."

                      Threat Modeling

                      Tanya first saw threat modeling when the CISO brought Tanya, newly on the security team, to a meeting with the business about what kept them up at night, and found the business worried about completely different things than a developer who just wants the app to stay up. The idea is to work out the threats to your system, then fix or protect against them, or accept the small ones. Formal frameworks like STRIDE and PASTA exist, but an informal conversation works as a warm-up, such as asking how you would hack your own app. Tanya tells of a friend with an IoT app whose view of the threats changed when Tanya asked about the users, and says it's "basically evil brainstorming, and the more point of views you have, the better."

                      Where to Start

                      Tanya suggests a half-hour to hour meeting with someone from the business, someone from tech and a security person, asking about confidentiality, integrity and availability: how sensitive the data is and where it's stored, what happens if something changes it, and what could knock it down and what you can tolerate, from a pacemaker to a neighborhood flower shop. Don't start with attack trees and a long formal process, which can scare people away. Mistakes include thinking threat modeling is the only thing needed, using a heavy process, and starting late: doing it at the end beats not at all, but it's "so much cheaper to find a design flaw really early than at the end." The OWASP application threat modeling wiki page is a good place to start.

                      Pushing Left, Not Shifting

                      If you draw the system development lifecycle left to right, from requirements and design to coding, testing and release, left is earlier. Shifting left, Tanya says, implies everyone's on board, while at previous workplaces Tanya and a friend from the Canadian government had to fight to start security earlier, so it was pushing. The blog series runs 14 posts because the fifth one was going to be 20 pages: a security activity at each stage, such as security requirements like HTTPS-only and key strength, secure design principles and threat modeling, secure coding and code review, static analysis and dynamic scanning. Tanya writes it as what Tanya wished someone had said two years earlier. Matty says we learn by teaching, and adds the swing-dance saying that "advanced dancers take beginner classes," so experts should read beginner material and just absorb it.

                      Serverless Security

                      Matty asks whether serverless is Amazon's problem. "No, serverless is not Amazon's problem." It's still an app, and functions appear and disappear, so a function that runs five minutes a week can be missed in testing, while malicious actors don't punch a clock. OWASP has a Top 10 for serverless risks, nearly the same as for web apps, and injection is still possible if a function talks to the operating system or a database. Keep an inventory, since "if you don't know you have them, how can you secure them?" Tanya has responded to an incident for an app no one knew the company had. Logging matters even when functions are fast, since without it there's nothing to investigate. Tanya says to log usernames and failed-login bursts, not social insurance numbers or dates of birth. Matty adds you can't log retroactively, and recalls an application error that simply said something has happened, and Tanya recalls apps that sent a daily ping to an inbox that everyone ignored. Matty calls it normalization of deviance.

                      What Developers and Ops Should Know

                      Tanya wishes developers knew the CIA triad, which isn't taught in school, and that the security team wants to help: "keep annoying us till you get what you need because that's our job is to help you." Tanya tells of a design that called double Base64 encoding encryption, and said another team had already built the real thing, but nobody knew who to ask. For searching, Tanya says "Whatever is at the top is the worst in regards to security every time," and recommends searching for OWASP cheat sheets for what you're trying to do, which surface the right answer. Matty adds that being good at searching has always been the secret, recalling using AltaVista in 1998.

                      For ops, Tanya says ops people get beaten up for unpatched systems though they work in slow waterfall settings, and that smaller, more frequent changes make emergency patches quicker. Security teams should buy ops people licenses and training for scanners like Nessus, and add container and VM scanning to pipelines. Tanya says assume breach and zero trust, recalling a network that drew zones on paper and was one flat network, and says a database should talk only to the app and its administrators, with the perimeter gone.

                      Mentoring Monday

                      Tanya mentors a few people and couldn't take on more, so started a #MentoringMonday tweet that thousands answered, now a weekly hashtag where people post what they want help with and others respond. Tanya retweets it, and women can also get a retweet from the WoSec account. It isn't only for infosec: Python, blockchain, project management and startups are welcome, since "Everyone is welcome." Tanya says to consider mentoring after two years in an industry, even if it's just naming the first book. Tanya recommends The DevOps Handbook, The Phoenix Project and Accelerate, and Matty says telling people to read those is most of Matty's job.

                      • Pushing Left, Like a Boss: Part 1
                      • OWASP Serverless Top 10 Project
                      • 49 min
                      • Catching Up With Steven Murawski

                        Trevor catches up with Steven Murawski at Microsoft Build 2019, the first appearance since Ignite. Steven is a cloud advocate at Microsoft focused on the operations side of DevOps, site reliability engineering and cloud-native operations, and the two worked together at Chef. They first met on an early episode about PowerShell Desired State Configuration, when Steven was at Stack Overflow. Trevor is still at Chef and has a workshop, a session and a keynote demo coming at ChefConf. The cold open is Steven joking about a movie scene where someone is taken out back and "you'd hear a bang."

                        Ignite the Tour and Learning From Failure

                        Ignite the Tour is Microsoft's 17-city, six-month, two-day event with Microsoft 365 content and Azure learning paths for migrating apps, running services in the cloud and hybrid operations. Steven's team ran the modern operations track: infrastructure as code, instrumenting applications, troubleshooting in the cloud when you can't crawl under the floor, scaling and global resilience. A favorite session, with input from Jason Hand, was about responding to and learning from failure. A colleague, David Blank-Edelman, coined the phrase "you cannot fire your way to reliable": if people fear for their jobs they'll minimize their part and stop sharing what happened, so the question becomes what about the system allowed the failure and how to engineer it out. For bad actors, Steven says to minimize the opportunity with a just culture, and recognizes that HR or the law may be involved.

                        Steven says a core SRE question is the appropriate level of reliability. Some software, like airplane systems or pacemakers, needs better than five nines, while others don't, and incidents can show that the impact on users is lower than expected, so a service might need four nines. Trevor's example is a coffee shop site that only matters from 9 to 5. Fail gracefully in dependent apps, or invest more where a service is critical.

                        Infrastructure as Code and Open Source Integrations

                        At Build, Steven's session was on infrastructure as code in a pipeline, focused on testing, because Steven likes code that does what it says. Steven has spent a lot of time on integrations with Ansible, Terraform, Jenkins and Spinnaker. People at Build see Azure and DevOps signs together and assume Azure DevOps is the only way to deploy to Azure, which isn't so. "We don't want you to have to change your toolchain just to be successful in Azure," whether it's Habitat, Chef, Ansible, Jenkins or Octopus Deploy. For people with nothing yet, Steven recommends Azure DevOps, but otherwise keep what works. Steven says a service company has to keep earning business: "If we can't make it easy and effective for you to consume our services, we're going to have a bad day," which Steven also liked at Chef's transition to services, unlike the old enterprise model of a pile of money up front.

                        PowerShell 7 and Windows Terminal

                        Steven is excited about PowerShell 7, the next open source drop, moving from PowerShell Core to just PowerShell, built on .NET Core 3, with expected compatibility of 70 to 90% with Windows PowerShell and a path forward from PowerShell 5.1 to bring back into Windows. Steven notes AWS bakes PowerShell into its Linux images, and describes PowerShell Summit the previous week, with a new on-ramp track and scholarships. The biggest thing for Steven is Windows Terminal, which can host WSL, command.exe and multiple side-by-side PowerShell versions, making it possible to test across versions without a box per release. It was due in June, and Steven said "I want it now." They also talk about Cortana as a framework automakers build assistants on.

                        Interest in SRE

                        Steven says a trend from the tour is interest in site reliability engineering, because operations people find its definitions more prescriptive than DevOps, which can feel fuzzy, like having a CI/CD pipeline. People ask whether SRE has to look like the Google book, and Steven says Microsoft is figuring out its own practice, and that service level indicators, objectives and error budgets give useful language for negotiating reliability. Monitoring also shifts: black-box monitoring of CPU and memory infers application behavior in a known environment, but in the cloud you need application performance metrics and the business drivers behind them. Trevor says the same question about DevOps having a uniform shape has the same answer: a core set of structures, and you figure out which fits.

                        Trevor chats with Steven Murawski of Microsoft about Azure DevOps, Windows Terminal and all the cool things from Microsoft Build 2019.

                        37 min
                      • Principal Engineering with Silvia Botros

                        Matty and Jessica Kerr talk with Silvia Botros of Twilio SendGrid about what it means to be a principal engineer. Silvia started as a Python developer at a since-gone CDN in New York, tripped over the database when it had issues and never left, and has been at SendGrid for about seven years, which grew from about 60 people to 500 and was acquired by Twilio about two months before. As of the Monday before recording, Silvia is a senior principal engineer, which mostly means "a lot more meetings." Silvia's org, SendGrid engineering, is about 130 engineers. The cold open is Silvia's line about being the org's archaeologist: "Table X, what does that do? I'll be like, let me tell you a story."

                        Tripping Onto Databases

                        Silvia says nobody grows up wanting to manage databases, since people trip on them and never come out, and Matty adds that nobody wants to be a sysadmin either, though an intern once said so and was hired. Silvia started with the DBA title at SendGrid, which changed to DBE as the role involved more code, and now works to expand beyond MySQL.

                        What a Principal Engineer Is

                        In Silvia's org, the role "is not like a senior, senior engineer." It is more strategic and business-oriented, built on influence without a manager's title or performance review authority. Silvia says a principal must be a force multiplier: Silvia's own shift over about a year and a half was from writing code to teaching others how to write it without trouble down the line, and in a large org principal engineers write design documents and help the team build, with mentoring the biggest part of the job. They can have a home area of the stack and still need to be T-shaped.

                        The skill Silvia calls most controversial is learning to talk to people other than engineers: product managers, finance and security. Principal engineers answer security and compliance questions about encryption and backups. Silvia says it would be a red flag for a principal not to understand what problem is being solved for customers.

                        Talking to Product

                        The process at SendGrid starts with a product canvas that lays out the problem and customer type, followed by solution validation with engineers, where principal engineers come in and explain, in English, what a request will cost, like a multi-region consistent database requiring Spanner and a large budget. Silvia dislikes the tech community's dismissiveness toward "just the product person," who is the voice of the customer. Silvia admits it didn't come naturally, having once been a DBA who got cranky at customers who called an API too often, until realizing that the company let them. Rate limits are an example: if you allow a behavior, you need to support it. Jessica calls it setting expectations. Silvia says "That's the biggest part of a principal engineer's job, is to make sure that what we're promising is what we're building."

                        Blueprints

                        Once product settles on what to solve, the delivery team writes a blueprint in a Google Doc covering what they're building, which helps onboarding and lets changes be explicit, with product able to see and comment on technical limits such as a service's SLA. An architecture team made of senior principal engineers reviews blueprints for one-way doors, decisions that can't be undone. It's a gate to production but not to proof-of-concept work. Silvia says the process can look waterfall-ish but tries to keep it fast, because customers build businesses on the product, and it shouldn't reach production by accident. Jessica says rewrites are appealing because it's the only time requirements are nearly complete, and Silvia adds that rewrites need a higher bar than Go being cool. Silvia notes that SLA math starts early: a service promising four nines in one region gives fewer nines across regions, and the more components, the lower the overall SLA.

                        Titles and Challenges

                        Asked about misuse, Silvia says "All over Silicon Valley," a title lottery, and "Staff engineer at Google does not equal principal engineer or architect at a company that's 18 months old." Jessica says it's about salary bands. The biggest challenge is the calendar, and finding the middle ground, since senior people become aware of other limits such as customer revenue, deadlines and security risks. Silvia's team motto is "strong opinions, loosely held," and Silvia used to be a no person earlier at SendGrid and earned flak for it. Jessica says the job is not to say no but "how do we get to yes?" Silvia also names the calendar as the best part, since it lets Silvia swap the DBA hat for a product or security one. Principal engineers partner with engineering managers, who are less in tune with the technical implementation.

                        The Org's Archaeologist

                        SendGrid hires principal engineers from outside, and onboarding is a team exercise, like Support Bootcamp, a four-day course where the support team teaches how to use every part of the product. Silvia, with the longest history, is one of the org's archaeologists who can explain what a table does. Jessica says to find such people and make friends with them. Silvia is now learning data stores beyond MySQL, and is a pragmatist: none will work all the time, and the question is whether somebody else found the sharp edges. Matty jokes that Silvia disrupts electronics, and Silvia tells of a hotel booking that failed twice and then said it was already booked, "I'm a living Jepsen," and a network flap at a Chicago data center the moment Silvia landed in Denver.

                        Advice

                        Silvia's advice for aspiring technical leaders: strong opinions loosely held, be an enabler, not the person who always says no, expect to spend a good chunk of time mentoring, and learn why things work the way they do, with healthy skepticism of new tools. Silvia is "very much of the Dan McKinley school of like, use boring tools to build cool things." Silvia recommends a talk by Tanya Reilly on glue work, and says "the internet is duct taped together with Bash."

                        • On being a principal engineer - blog post by Silvia

                        • Silvia's talks

                        • Image credit

                          Community
                          • Velocity San Jose June 10-13 2019 - discount code "ADO2019" gives 20% off for Gold, Silver, and Bronze passes.

                          • For any devopsdays, try the discount code ADO2019!

                          • CFP for devopsdays Chicago: open until May 3rd

                          • Checkouts
                            • Silvia: the Beyoncé movie came out on Netflix: Homecoming

                            • Jessica: Do something outside! It’s spring!

                            • Matty: Stocksy - for affordable stock imagery that benefits the artists and Super Team Deluxe for great pins and stuff

                            • 52 min
                            • Making DevOps Magic with Arup Chakrabarti

                              Matty talks with Arup Chakrabarti, a director of engineering at PagerDuty who previously worked at Amazon and Netflix, about guiding DevOps transformations. Both work at PagerDuty, and Arup says a draw of the job was helping customers reach the business results the large consumer companies got. The cold open is Matty: "The real world will destroy all your plans."

                              Start Small, Find Allies

                              Arup's step zero is to start small, with one team, project or codebase instead of a 10,000-person department, then find champions, meaning people who reply to your emails about change with enthusiasm, or build them by giving context. Arup says the journey gets lonely, and allies make it easier. Matty compares it to making a movie, since there are days you hate it. Arup adds that leaders should expect some short-term negative business impact, such as more incidents, before long-term gains, and that conviction tends to spread.

                              Metrics and Context

                              Arup likes number of deploys as a metric, because it is a proxy for operational maturity, including incident response and end-to-end ownership, and suggests moving from yearly to quarterly, not straight to continuous deployment. After that comes SLAs, SLOs and SLIs, and the Google SRE book. Arup describes a previous company that tracked revenue no longer lost to downtime, reviewed weekly with the question of whether the team had done something to change it. Matty says "you can't know if you're moving the needle if you don't have a needle," and stresses there's no magic number of deploys, so Amazon's frequency isn't a target. Arup says there is a lot of DevOps FOMO, and "you're not PagerDuty. You're not any of the companies that you saw at a conference," and quoting another company's way alienates stakeholders.

                              Matty adds that stage talks are success stories, because speakers feel better telling them and PR departments block failure stories. Arup says a bank has constraints but can use its customers' priority of an accurate ledger as an advantage. Matty recalls listening to sales calls at Apartments.com after an office move, learning the value of a lead, and says to learn how your company makes money. Arup says to talk to finance and product managers. Arup's examples of a right metric: Amazon's order graph, Netflix's stream starts per minute, and three at PagerDuty, endpoint availability, time from event to incident to notification, and web experience, which took years to settle. Metrics are proxies, so halving incidents doesn't double the customer experience, "but we know we've made it better."

                              When Metrics Backfire

                              Matty says to question metrics, asking why, and to avoid targets like "not slower than last month." Arup says to set metrics early, then "question the crap out of it" a month or two in, and tells a story of measuring availability as the percentage of 200 responses, when an engineer served every 500 from the load balancer as an empty 200, and the team congratulated itself until the manager asked what happened. Matty cites Andrew Clay Shafer on people working to metrics to the organization's detriment, and Jez Humble's story of a goal of one test per sprint that produced assert-equals-true tests. Matty says "There's a difference between being committed and being compliant," and that people usually just lack the why. Arup compares intent of a metric to the intent of a law. Arup tracks median pull request duration but sets no goal on it, so it stays a temperature reading, though an engineer noticed it could be gamed. Matty says don't litigate severity in an incident call, and that when teams are measured on counts of Sev 1s, the metric becomes "mean time to innocence."

                              Anti-Patterns

                              Arup says a common mistake is expecting a transformation in months: "if it took you a decade to get into a problem, it's going to take at least a year to get out of it," and you're never done. Matty adds the urge to plan everything first, recalling a large insurance company that took six months to plan a first change with Chef. Arup says rigid plans and assuming no risk are dangerous, and starting small accepts a bit of risk that de-risks later.

                              Treat It as an Experiment

                              Matty draws the parallel to chaos engineering: a hypothesis, a limited blast radius, known measures and a learning experience, though the word experiment may unsettle stakeholders. Arup likes the word because it implies humility, and tells stakeholders Arup will be first to acknowledge a failure. Arup describes PagerDuty's introduction of Chef about seven years earlier as an experiment at the lower levels of the organization that over about a year and a half became the way. Another customer of around 5,000 engineers feared engineers would quit if everyone went on call, and so tracked the number who quit and slowed the rollout if it rose, and the number went down as they invested in explaining why. Matty warns about mistaking correlation for causation, such as an acquisition happening alongside.

                              Wrapping Up

                              Arup's tactic for any leader who can't say how the business makes money is to go talk to the CFO's finance team, who are transparent, and jokes "Can't boil everything down to a single shell script, unfortunately." Matty says to find a buddy if you're not good at selling change, and to make the tent big, including testers, product and FinOps, since DevOps is unfortunately named. Arup says DevOps means pulling in whatever stakeholders you need and owning the problem, not throwing it over the fence, including to product management.

                              • Image credit: photo by GotCredit
                              • Community
                                • Velocity San Jose June 10-13 2019 - discount code "ADO2019" gives 20% off for Gold, Silver, and Bronze passes.

                                • For any devopsdays, try the discount code ADO2019!

                                • 50 min

                                About Arrested DevOps

                                From the publisher's feed

                                Arrested DevOps is the podcast that helps you achieve understanding, develop good practices, and operate your team and organization for maximum DevOps awesomeness.