
Sign up to save your podcasts
Or


Matty talks with Laura Santamaria, a developer advocate at LogDNA who comes from a self-taught background, about making DevOps friendlier to people who are new to it. Laura has a science degree, worked in teaching and education at a science museum, and then moved into programming, where Laura owned a first production system as a junior engineer. Matty notes that when the show started in 2013 it was meant to be a beginner podcast, for "the people where your boss read about DevOps in the in-flight magazine and came to you and said, hey, we need some DevOps now," and is amazed it took this long to do an episode on the topic. The cold open is Laura's line that "this isn't some type of snake oil."
Laura says the definition of DevOps is hard to nail down, and "you can't buy DevOps," since it's a cultural shift. Getting into that culture takes history and hands-on experience of things not working, but you need to be part of that world to get the experience, a chicken and egg problem. Matty says context matters: people who've "been in the shit" see why it's better, while newcomers may not see why to push through the hard parts, or may wonder why it's a whole movement at all. Laura is one of the people for whom it was common sense, having started on a system built by a team focused on DevOps, and forgets that others find it unfamiliar, like being told dev does this, hands it to ops, and marketing and sales are elsewhere.
Matty says people worked the old way because those structures solved the problems of their time, not because they were dumb. Matty recalls PagerDuty new-hire orientation, where newer colleagues couldn't see why managing notifications and on-call schedules had been a problem, while Matty remembered spreadsheets and email aliases. Matty also points out that some DevOps beginners are senior technologists who are new to this approach.
Laura says the biggest thing for any newcomer is understanding "you're allowed to ask questions." People who've been around in a waterfall culture may not feel it's a safe place to ask why they're being asked to do something, and Laura says a good devopsdays provides that space through open spaces and people who are open to "ask me anything." The hard part is getting people there to begin with.
Matty adds that veterans are jaded because advocates used to be the digital darlings like Netflix and Facebook, and people replied "that's great, but we're a bank." Matty says DevOps Enterprise Summit was among the most powerful things to happen to the movement, since large enterprises and government organizations started talking about their work, and people could say "now you're talking my language." Once you know something can be done, you can probably do it, and "if you don't know that it can be done, it can seem insurmountable."
Laura suggests offering help directly: Laura answers questions in Slack communities and gets DMs from people who are stuck, which models helpfulness. Matty adds that you have to meet people where they are, since not everybody goes to conferences, meetups or Twitter.
For a team welcoming someone new, Laura says to identify the people with the patience to answer the same question a few more times, since not everyone has it, match the person's style with documentation or a person, and set up sandboxes so they can learn by doing without taking down something important. Laura says "you're just going to figure it out as you go" is too intimidating.
Matty says some of this way of working inherently feels wrong to newcomers, like giving devs access to prod, so encourage them to say why they're uncomfortable, since they may raise things the team has stopped questioning. Matty recommends pairing and shadowing, and a living "rules of engagement" document reviewed a couple of times a year on how the team communicates and talks about problems, as a lot of it is undocumented knowledge. Laura likes having each new person update the docs, like being the scribe on an incident, and calls the norms guardrails that can be repaired if they head toward a cliff. Matty cautions that you shouldn't measure success by how many changes were made, but by validating and updating as needed, and says even the author of a runbook should follow it. Laura, who has been a technical writer, says to audit docs every time, because the arrow may now be on the left of the button.
Laura says joining any group is intimidating, since you don't know who's safe to ask, and that there's so much going on it's hard to know where to go. Laura was lucky to fall in with a meetup and volunteer at devopsdays Austin, and didn't learn about the community Slack for years. Laura also says that for people with experience, "all your war stories" can be intimidating for people without them.
Matty suggests sharing experience as "thank goodness this isn't true anymore," not "you sweet summer child," and asking what still sucks. Laura's own war stories are about communication, and Laura says marketing, support and sales have the same stories, including the morning after when the emails come in. Matty adds that some exclusivity can be fine in a space such as an open space on AS/400 admin war stories, not as a requirement for the bigger conversation, and Laura calls this a community of practice, like birds of a feather.
Matty's tactical point is that tutorials should tell a story of why: "Why does Kubernetes matter?" and "Why did I need an S3 bucket?" Make them journeys, since context is where it fits "into the big smushy world." Laura's closing idea is that DevOps is about zooming in and out, like a camera lens that catches the butterfly on the flower and the whole picture.
Laura will be at Spring Live on March 19th, runs A Minute on the Mic, and is @nimbanatus on Twitter. Matty will speak at FailoverConf on April 21st.
Check out A Minute on the Mic for bite-sized videos from experts on various topics!
Matty and Jeff Smith, in Jeff's first episode as a co-host, talk with Patrick Debois about Patrick's new role and journey. Patrick recently became Director of DevOps Relations at Snyk, a DevRel role that's new to Patrick, aimed at talking about security from an ops and DevOps perspective. The word "relations" matters to Patrick because it makes the job less about technology and more about people collaborating and the cultural side. The cold open is Patrick's line from the 2010 conference story below: "In all due respect, I call bullshit."
Jeff asks what gave Patrick the nerve to change an industry. Patrick says "I didn't have the idea," and never takes credit for it. Patrick put a lot of effort into promoting the ideas, but they weren't Patrick's, and compares it to research, where you find little nuggets and put the pieces together. It stuck because the first devopsdays got people so excited that they took it further, to the US and Australia, and Patrick says the timing was about "amplifying the zeitgeist." Patrick admits to organizing the first events because it would be cool to travel to country X, finding random people on Twitter, and says the drive was learning, which "might sound selfish."
Matty mentions the 10th anniversary of devopsdays in Ghent and the Day Zero for organizers, which had more people than the first devopsdays. Patrick calls it surreal, and is struck by the cooperation among random individuals on their free time, helping people on a budget fly to Belgium, and by growth in places like Brazil and China. It "sometimes does feel a little bit like a religion," and Patrick rants on Twitter about the echo chamber in dark moods.
The big thing Patrick learned from the global community is that tools and human relationships go hand in hand, though we don't want to believe it. Patrick points to the irony of automation and promise theory, which Patrick learned about from Mark Burgess, plus safety and resilience. Jeff asks why safety conversations thrive on Twitter but not inside organizations, where people are still being asked to do blameless postmortems. Patrick says there's a lag, and that change takes time, including time to recognize you need to change: "it's really hard to sometimes see yourself honestly in the mirror." Matty says devopsdays has always had an underlying human factors approach, so resilience engineering fits in naturally.
Patrick recommends a forthcoming IT Revolution book by Jeffrey Fredericks and Douglas Squirrel about honest conversations, where you say you'll go along with tool A while thinking something else, and being transparent about what you mean is part of safety at work.
Jeff asks how teams evaluate what they're capable of delivering. Patrick says the first step is being aware you're not good at something. Patrick first thought more talking meant healthier collaboration, then read that it's not the number of interactions that makes a better party. Patrick's paradox is that we now use services and don't talk to them, so is that the end of DevOps? Patrick's answer is that you don't need to communicate if the other side helps your goals, and "if you don't hear too much of the friction points, then things are good." Matty adds Dunning-Kruger and a point about not knowing how far the spectrum goes, and says sharing at devopsdays and DevOps Enterprise Summit resets where people think they are.
Jeff says it takes vulnerability to be the first to say you don't know what an acronym means. Matty tells of a game with a fellow systems engineer, where one texted the other a made-up phrase on a two-way pager and the other used it in a meeting as if it were a saying, and nobody ever questioned it. Patrick recalls a 2010 conference where Stephen Nelson-Smith, "this impeccable Englishman," stood at the microphone during a panel and said, "in all due respect, I call bullshit." Patrick says Patrick was shy, and learned to speak up because people kept asking for an opinion. Patrick's tip from someone else: you don't have to talk like you know everything, just say what you found and ask for another opinion.
Jeff asks about people in an organization who aren't bought in. Patrick wouldn't go in saying it's what everybody is doing, but would try to understand the person's problem, because bad past experiences make trust hard, and says you first learn the history and context, like a firewall auditor who flags open ports that the business needed. Patrick prefers supporting people with little successes to being on the barricades, and says if you can't change beliefs "you have 2 options. Either you leave or they leave," while wanting diversity of opinions too. Change is baby steps: find champions, book successes, and question yourself either way. Jeff adds that organizations have preferences, and that the purpose of the business isn't the technology.
Patrick tells of spending about 4 years at a mobile and video streaming company, getting the tech straight in a year and improving sales, marketing and legal, and realizing "we're not making any money, so there is no business fit." The bottleneck "just moved, moved, moved." Jeff says the win you celebrate this week is next week's bottleneck.
On project management, Patrick mentions the No Estimates book and says to treat plans as a projection, play the milestone game while asking for a stake in the ground on features, time or resources, and build confidence by delivering smaller chunks, talking to project managers about their real goals.
Jeff asks whether it's fine to skip a technology, like jumping from the data center straight to Kubernetes, compared with the leapfrog to mobile. Patrick notes Jeff is already biased, and says to ask what you're trying to achieve instead of which tool to use. At a recent meetup, asked who isn't using Kubernetes, Patrick was the only one to raise a hand, because Patrick felt no friction in what they were using. Patrick says being mature is picking the right tool at the right time, same with serverless, and squads at Spotify, which Patrick says Spotify doesn't use anymore, and to avoid "resume-driven development or operations." Jeff jokes the episode title is that Patrick Debois doesn't need to run Kubernetes.
For a new domain, Patrick starts with blog posts to find the keywords, then buys several books for different perspectives, since they share about 70%, then follows blogs, looks at slides on SlideShare, and finally follows people on Twitter. Patrick is on Twitter and will speak at QCon London, DevOps Talks Melbourne, DevSecCon Sydney and the new O'Reilly conference in San Francisco. Matty will be at devopsdays New York March 3rd and 4th and runs a workshop on blameless postmortems at SREcon Americas West, and mentions devopsdays Chicago's CFP, in its seventh year.
The panel discusses the origins of the book Team Topologies. The project started with a blog post.
Matthew: “Back in 2013, I actually wrote a blog post in my personal blog. I actually wrote it in a rage.”
In 2015, Manuel joined the team to help expand on the ideas from that blog post and create Devops Toplogies.
Manuel: “What the hell are you calling a DevOps team? DevOps is not about creating a new team called DevOps.”
The panel discusses the impact of DevOps Topologies and some of the companies that have used it, including Netflix and Conde Nast.
Matthew and Manuel explain how the project has evolved over time as DevOps Topologies was being deployed in the real world.
Matthew: “It’s not just a set of patterns or templates. We wanted to provide an organizational capability for detecting when things have changed and have gone wrong.”
The panel discusses Conway’s Law and its implications for DevOps.
Manuel: “Teams are the means of delivering value.”
The panel discusses the importance of flow in both living systems and organizations.
Manuel: “It’s a more experiment driven approach where we have this goal or this need we need to meet and then allowing the teams to find the right solution.”
Jessica: “In modern systems, the flow is of changes to the flow of the product. It’s a very different level of work.”
The panel discusses the need for different team configurations that are constantly evolving.
Matthew: “It seems quite important to understand different kinds of dynamics in the organization at different times.”
The panel discusses the three team interaction modes laid out in the book: collaboration, as-a-service, and facilitation.
The panel discusses the four team topologies in the book: value stream aligned teams, enabling teams, platform teams, and complicated subsystem teams.
Jessica: “The limitation of a team is cognitive load. It’s not resources, it’s not pizza.”
Manuel: “Although pizza is very appealing.”
Manuel discusses Dunbar’s Number and how that concept can put useful constraints on teams.
Buy the book!
Episode images by Jessica Kerr. Show notes by Tyler Wilson.
Matty talks with Jono Bacon, author of the book People Powered: How Communities Can Supercharge Your Business, Brand, and Teams, and a consultant, author, advisor and speaker on community and collaboration. Jono got into Linux in 1998, which is where Jono discovered communities, and has worked at Canonical, GitHub and XPRIZE. Jono's view is that healthy communities improve the human condition, and that people "don't want the newsletter anymore" but a relationship with the companies and projects they follow. The cold open is Jono on lawyers: "What's the best way of reducing risk? Doing nothing."
The book opens with an email Jono got about six months into working at Canonical, from a kid in rural Africa with no computer at home. The kid earned money from chores in the village, walked about two hours to a town, bought an hour of internet at a cafe to contribute to Ubuntu, and walked two hours back. Jono says it showed what's possible when someone feels part of a bigger mission, and notes that Ubuntu is an ancient African word meaning humanity towards others.
Jono defines a community as a group of people interconnected around a mission or an ethos, and says social media can play a role but isn't one by itself. Jono breaks communities into three models. Consumers share a common interest, such as Star Trek fans or Kubernetes users, without much influence on the core thing. Champions go the extra mile with content, videos, blog posts and events, which add to a stockpile so the community gets more valuable. Collaborators come together to build something, either inner collaborators on exactly the same project, as in open source, or outer collaborators building things on top of your platform, like WordPress plugins or npm modules.
Across those models, personas have different needs. Jono's example is a company building a community for business decision makers who will never go to a forum and communicate by phone and email, alongside technical implementers who are used to forums, Slack and Stack Exchange. Defining the personas, including where they hang out, what motivates them and what they fear, makes the other decisions easier. Jono says you can't attract everyone at once, so you have to decide who is most critical.
Jono splits communication into structured, such as a GitHub or GitLab issue or a Stack Overflow question, where there's no "how was your weekend," and unstructured, which is everything else. Unstructured platforms divide into short-term and long-term memory, a term Jono borrowed from Jeff Atwood. Slack and Mattermost are gratifying but transient, and "anyone who claims that they can do this is lying to you" about finding old discussions, so the same question gets answered repeatedly. Forums such as Discourse are easier to reuse, and Jono likes the Solved plugin, which lets you mark the 16th post as the answer and gets found on Google. Stack Exchange and Askbot, in contrast, "build value, but you don't build community." Matty says that the conversation around an answer has value too.
Jono agrees you can't create advocates, but you can facilitate them. The project has to be interesting, and people are motivated by dreaming big, so build a mission around what you're doing. Jono describes a member journey of casual members, then regulars, then core members, encouraged by incentives and mentoring. When someone's name keeps showing up, ask them to do things, without pressure. Jono's example is content: plan the tutorials, demos and events you'd like, then ask community members to write them with editing, guidance and promotion. "You just got to ask."
Matty wonders about legal and PR worries over outsiders contributing. Jono says it varies, startups are less anxious, and "lawyers are not incentivized to let things through. They're incentivized to reduce risk." Companies should tell the community, "you're not giving any control," but treating people more like equals and setting clear expectations about where the line is.
Jono says a common mistake is tacitly agreeing to everything for fear of upsetting the community, when it's better to say where the line is. Jono puts it as "I'd rather piss you off now than piss you off in 6 months." Companies that do well build inclusive environments, with open issue tracking, open code and regular meetings, where engineers, product people and the CEO are members of the community too. Jono names HashiCorp, Docker, Microsoft, Red Hat and Mattermost, with Red Hat praised for training non-open-source staff in open source nuance. Matty adds a saying about the first rule of open source, that "no is temporary, yes is forever."
Jono says to slice metrics by persona and avoid "data fetishism" dashboards with hundreds of graphs. For contributors, tangible metrics include pull requests submitted, time to first response, time to review and merge, and issue lifespans, and intangible ones, like whether people are happy and learning, are best tracked by community managers. Track the action and also its validation, such as a pull request being merged. Avoid counting repetition, like a forum ranking after 500 posts, which rewards repetitive posting. Pick 5 to 10 metrics and look at patterns on a regular cadence, asking whether a slowdown is summer or a new release.
For keeping up with personas unlike you, Jono's advice is to be organized, breaking a big list of tactics down by quarter, and to be intentional, talking regularly to the people involved. Jono would put a recurring open-ended meeting on the calendar with translators, for example, since "all of the answers to how to do this kind of stuff well live in your audience's brains." Surveys give some data, but a call gets more valuable answers, especially if you ask where you can make a bigger difference. Jono's last advice is to start simple and track a limited set of metrics for a limited set of personas.
Switched from devrel to product. Helped with Helm 3 launch!!! Traveled less. Dragged Joe to a personal trainer because if he’s going maybe I will too.
Lots of speaking (spoke at 24 conferences in 2019). Moved back to Chicago. Went to lots of devopsdays. Got even more interested in resilience engineering and realized I don’t know shit.
More travel! A lot of customer work, focus on delivering solutions, now moving officially into product. I also got really into shuffleboard and DND
Consulting!
Matty records at All Things Open in Raleigh, North Carolina, on the eve of the conference and Matty's first time there, with speaker Ali Spittel. Ali started as a software engineer and moved into teaching code, as an instructional lead and distinguished faculty member at General Assembly, a coding bootcamp, teaching roughly 30 people at a time from zero to professional software engineer. Ali has also found that blogging and podcasting reach a bigger audience than a classroom. Ali is giving a workshop on blogging to gain visibility for open source software, and a talk on teaching code. The cold open is Ali's line: "You have a unique perspective, no matter what your perspective is."
Ali says that if you're going to advance in a development career you'll end up mentoring, leading a team or moving into management, so knowing how to onboard people and help them learn concepts, not just fix one bug, pays off. Teaching also makes you a better programmer, because people ask questions you'd never think of, and "if you can teach something well, you really know it." Ali took education classes in college, would have been an education minor if they hadn't left to be a software engineer, and got more from them than from computer science classes.
Matty says teaching Chef Fundamentals workshops at customers for two days over and over taught Matty things about the product they hadn't known. Teaching forces you to think about the why, and the process stuff sticks more with some background behind it. Ali's key teaching concept is linking: take what somebody already knows and link it to a new piece of knowledge, because teaching in isolation is "impossible to learn," which is why a first programming language is harder than a second.
Matty asks how to do linking in a big room or a recorded course. Ali's answer is to subtly reteach things several times with different formats, such as diagrams, videos, a code-along and a traditional lecture, and to use different examples for different people. Ali notes that Star Wars tutorials appeal to some, but "my students freak out when I do Britney Spears fans apps." When teaching variables and functions Ali links them back to algebra class.
Matty adds that as learners we think we know more than we do, recalling the swing dancing saying that beginner dancers take intermediate classes, intermediate dancers take advanced classes and advanced dancers take beginner classes. Ali agrees from experience: after years as a software engineer who jumped straight into React, teaching at a bootcamp meant relearning foundational JavaScript from the ground up.
Ali says everyone has a unique perspective, and it can be terrifying to publish when someone has done it before. Ali wrote a React tutorial well after the initial burst of React tutorials, sure nobody would read it, and it became the most read post. Matty tells of a SharePoint fix posted on a blog mostly as self-documentation, which an MSDN forum answer then sent thousands of views a day. Matty also says some posts have an intended audience of "future Matt."
Ali has a talk called Yes, You Should Write That Blog Post, about writing for three people: past self, present self, since teaching cements learning, and future self. Ali returns every year to a post on filling out calls for papers. Matty's popular post was on getting a Powerline font working on a Mac in Visual Studio Code, and Ali's was about how they set up computers, "nobody wants to read that. Well, turns out they do."
Ali's advice is to make it skimmable, since "people don't want to read an essay," with headers, images and lists, and to add the why, what problem this solves, in the introduction. Also appeal to multiple learning styles with code demos, graphics and other formats such as podcasting and video. Matty says lists get people in, and then they find your deeper posts. On talks, Matty calls the why versus how balance "eat your vegetables" and likes single-track conferences for that, and Ali says conference talks aren't the best format for education because people aren't hands-on, so the best thing is to get people excited enough to learn afterward. Matty says attendees asking for more technical content usually mean practical examples.
Ali says the two main channels are social media and search engine optimization. On social media, consistency and being a real person matter. Matty says you don't have to share everything or engage with everyone, "You don't owe anybody a conversation," but to be successful "have some conversations." Ali adds that LinkedIn posts live longer than tweets. For SEO, Ali says to add a couple of keywords but not to keyword stuff, and backlinks come from making good content and relationships, not spamming. Matty adds that when asking for a favor, "don't give 'em more work to do."
Ali recommends cross-posting to a platform with a built-in audience such as dev.to, using a canonical URL so Google knows it isn't duplicate content. Matty suggests the company engineering blog, which is usually "dying for content," and says that cross-posting with a canonical URL helps everyone and is good for recruiting. Matty also says repurposing talks as blog posts and the reverse works well. Ali says making it part of the workday is one strategy for finding time.
Ali says the number one question in blogging workshops is how to handle people online, and that when criticism gets personal it comes from "a place of hurt." Ali expected the worst from day one as a young woman on the internet, but it didn't happen for about a year, until a post hit the front page of Reddit, so it probably won't be as bad as you think. Ali's own strategy is to write out a response, screenshot it to a couple of close friends, and delete it. Matty says that if you're someone who doesn't have to deal with it, you can be part of a vent network: say that's really shitty, and that you're sorry, and be done. Ali adds not to minimize it, because "nobody should have to deal with that."
Ali suggests writing the post you were Googling for and couldn't find, writing your own story, or writing about what you want to learn, and Ali's first blog was learning a new technology every week, building an app and writing about it as a beginner. Matty tells of Annie Hedgepeth, who blogged while learning Inspec, and for a long time the project's documentation was that blog. Ali likes ConvertKit's slogan "teach everything that you know" and likes complete beginner guides. Ali's biggest tip is to write down ideas when they come, while walking a dog or Googling, and not when you're looking to write. Matty ends by asking for ideas for 2020 talks on Twitter at @MattStratton and pointing to devopsdays.org/speaking for open CFPs.
Bridget talks with Kelsey Hightower, who describes themselves as "a minimalist" who enjoys "learning in public and helping other people do the same." The conversation grows out of a Twitter post where Kelsey said they wanted to go on podcasts and talk about where Kubernetes is going. Kelsey's answer is that it's "not the future, it's the now."
Kelsey compares Kubernetes to the 56K modem: some people liked the dial-up sound because it meant they were getting on the internet, and people now look at their nodes and clusters and feel like they're doing computing. The internet got interesting when that went away, with DSL and wireless routers, and now "the internet is now just a thing." Kelsey says "Things tend to get better when they disappear," and that we're in the 56K modem era of Kubernetes, and hiding it will let more people use it without learning to manage it.
Bridget raises the fast-moving release cadence, and someone stuck on 1.12 or with no Kubernetes yet. Kelsey says that's a good place to be, because "if you don't have this problem, you don't really need this solution quite yet." Linux went through the same transition, from rolling your own distributions to Red Hat and Canonical, and Android users benefit from Linux without touching the kernel. Kelsey says it's early: Kubernetes is only 6 years old, VMs still work, and some people will skip containers and go straight to serverless, but Kubernetes-style APIs are resonating, and some people will use parts of it without ever being a cluster administrator. Kelsey describes GKE, AKS, Fargate and K3s as "checkpoints" in an ever-moving project, and says Kelsey is now a consumer of Kubernetes who goes to the checkpoints and uses it as is.
Asked what people can still learn from Kubernetes the Hard Way, Kelsey recalls that early on there were no docs, and learning how to install it came before the first line of code or PR, which revealed what the scheduler did and what went where. Kelsey's argument is "you can't really fix a system that you don't know how it works." The guide goes step by step with no scripts, so people see what the kubelet does, how it connects to the API server, and where the certificate goes, and gain the foundation to troubleshoot and debug.
For production, Kelsey says there are many layers, and security is number one, since performance can be tuned later but "once that security hole is too big, it's a little too late to go rewind the clock on a breach." Most clusters come out of the box with flexibility and not security, so you can run random images and run things as root, and Kelsey compares it to SELinux, which every system admin turns off. Kelsey says 30% of Kubernetes the Hard Way comes from security feedback, which is why it generates certificates for each component and encrypts secrets in etcd, and why it leaves the dashboard out. Bridget says they show the dashboard in workshops with a giant disclaimer about cryptocurrency miners. Kelsey is encouraged that KubeCon talks and docs now say to do exactly 4 or 5 things, raising the security profile for many people at once. Bridget mentions Ian Coldwater's keynote and Gareth Rushgrove's work with Open Policy Agent.
Kelsey uses CDNs as the model. People once glorified FTP and then SFTP, and thought FTP would just become more robust, but CDNs took the problem of getting files close to people and made it disappear as the complexity grew and the number of people who understood it shrank. Compute is slower because so many people believe they understand the compute problem, so a new person invents a new platform every couple of years. Serverless says there are about 80% of compute use cases that are understood and never need to be built again, but mainframes and VMs don't go away either. Kelsey says "I really look at this as we're going to have multiple things in parallel," and that if starting from scratch, "I would probably try to go as high as I could and focus on building my app and the product before going to play infrastructure again."
On trust and lock-in, Kelsey says nobody digs up a wire to connect to the backbone of the internet, and compute needs providers that earn trust with open interfaces. On compliance, Kelsey says some banks are 100% online and see a single building with a door as too much risk, so it's "different degrees of understanding the risk."
Kelsey asks engineers and executives to stop using the word legacy, since it has a derogatory context, and says "classic infrastructure" instead: "It's the stuff that actually worked cutting everyone's paycheck." Bridget calls it "the place where all the customers and money are." Kelsey asks what problems they have now, such as service discovery or scripts and large on-call rotations for failing over, which is the opportunity, since a part of Kubernetes solves that specifically. They may not need all of Kubernetes, and Kelsey asks them to have a good reason why: if 15 people maintain a scheduler that looks like Kubernetes, those people could work on another problem, or half of them could contribute to Kubernetes.
Kelsey likes to spend time on-site asking what's on the backlog, often observability and security, and says even if Kubernetes automated you out of a job, "which it won't," it would give you time back. For people deciding whether to adopt, Kelsey says the world will continue to move with or without you, so put the tool through its paces, and if the answer is no, write an internal doc on the reasons and revisit if they get solved. That's "just engineering."
Kelsey has worked at Puppet Labs, used Ansible and contributed to Ansible and Terraform, and describes the last 10 to 15 years as an attempt at infrastructure as code, with DSLs, for loops and then the problems of any codebase. Kubernetes tries something slightly different: "No more infrastructure as code. Now, we're doing infrastructure as data." The logic and state machine live in the controller, which some people call operators, and on the front end you restrict yourself to data, which is YAML, even though Kubernetes itself only supports JSON and protocol buffers. The drawback is a lot of YAML, but like assembly language, any language can compile down to it. Helm can run as a preprocessor, pipe to Kustomize to patch, and go through an admission controller, which is "the dream come true" of describing infrastructure with a type system and interchanging tools.
Bridget asks whether the flexibility makes the complexity unapproachable. Kelsey says Kubernetes is "formalizing the complexity," so it can be seen in one place, and people deal with it at different levels. Someone managing the environment can create a custom resource definition that says deploy my app across 50 countries, and control loops do the heavy lifting. Kelsey says that if you just want to deploy containers, "writing CRDs and operators is the equivalent of writing kernel modules," a job for the people who need to extend the system and not for most people, who need only declare a load balancer, a certificate, a DNS name and a container.
Bridget says picking the technology first is resume-driven development. Kelsey says it's hard when you don't know what question to ask and every deploy goes wrong for 10 years in a row, and then you go to KubeCon and see someone deploying all over the world, and think you need some Kubernetes, which Kelsey admits to being partly responsible for. Bridget jokes, "Ask your doctor if Kubernetes is right for you."
Kelsey's story is teaching their daughter to make a GeoCities-style web page in a text editor, viewed in Chrome. When the daughter asked why a friend couldn't see it at 127.0.0.1, Kelsey used Firebase, and "she said Firebase deploy and it spit out a URL." Kelsey calls that the end game: people with an idea and finished code who want to see it come to life, and any platform that gets closer to that moment is exciting. Kubernetes will evolve that way from the ground up and serverless will work its way down to support other workloads. Bridget adds that we should remember we're building things that produce actual value, and Kelsey agrees, adding that the middle layers matter and aren't the end game. Kelsey is on Twitter, DMs open, and on YouTube, with meetups announced about two weeks ahead.
Bridget chats with Kelsey Hightower about Kubernetes and the future.
If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf
Bridget talks with Ian Coldwater, recorded live at a meetup, about their KubeCon North America keynote and Kubernetes security. The cold open is Ian's line that "Attackers actually generally are unconcerned with whether or not you have your compliance boxes checked."
Bridget asks if Kubernetes security is possible. Ian says yes, but Kubernetes is not secure by default, so if you assume it's secure out of the box you may be unpleasantly surprised. Asked for the over-under on new CVEs between finishing the keynote and giving it, Ian says they don't personally know of anything anybody is sitting on, so the number is completely unknown. For learning, Ian points to the engineering blogs from Aqua Security, StackRox and Twistlock, the Kubernetes security announce list, the Kubernetes documentation, past KubeCon talks, and people on Twitter. Ian acknowledges it's a pile of stuff, because Kubernetes has a lot of moving parts, but "you can do it."
Ian's top item is admission control, "the biggest thing that you can do to stop attackers from being able to compromise your cluster." It's a set of policies that dictate what privileges get run and who is allowed access, and pod security policies are part of it. The user experience could use work, with pod security policies particularly notorious, but it's worth learning. Ian's second item is being careful about what's exposed to the internet, and suggests looking at Shodan, which indexes every 24 hours, so the idea of being too small to be a target isn't true. If SSH ports are exposed, use SSH bastions.
For developers who don't run the cluster, Ian says one way in is supply chain attacks. Software has a supply chain, and Ian uses npm as the example, where a Node module has dependencies that have dependencies. Libraries and container registries are all potential vectors, so know what you're running and keep it up to date. Ian adds that people are a vector too, including yourself: developers know they're smart and may be more vulnerable to phishing than marketing or sales, since they're confident and may "pay less attention to those trainings by the numbers," and they often have more access, so they make exciting targets.
On pinning versus staying current, Ian says containers done correctly make patching much easier, since you can spin down a container and spin up an updated one. In Kubernetes, have a plan for upgrading, and especially for upgrading in place, since newer security features may not come into an existing cluster to avoid breaking changes.
Bridget calls back to the keynote line. Ian explains that people who aren't attackers think in lists: a list of compliance boxes, or sprint points, and "did we get them all?" They aren't necessarily looking at the layout of what's running and how things connect. Attackers want to get in, find out what's there, how it talks to other things and whether anything is vulnerable, to get "a lay of the land." If you don't know the connections between your resources but the attackers do, "that's a disadvantage for you." For building the graph, Ian recommends a VMware tool (the transcript garbles its name) that's useful both to administrators and to pen testers, since you can run it locally with the privileges you have. Ian has heard of cluster visualization resources in Visual Studio Code but hasn't tried them. Bridget notes the Kubernetes extension for VS Code is open source and plugs colleagues who work on it.
Bridget asks why Ian is "suspiciously positive for a security person." Ian says security people are notorious for being the team of no, and that it doesn't help relationships, since people who don't like you won't talk to or listen to you, and doesn't help security either, because developers and operators have their ears to the ground. Ian has worked in DevOps and knows what sprints are like, thinks most people mean well, and says "being a negative jerk doesn't really seem like it helps with that, so I just don't do it."
For people who want to become pen testers or red team members, Ian recommends capture the flag games, in which the flags are on servers you're sanctioned to compromise, and mentions overthewire.org as a beginner-friendly site. Ian has put on CTFs internally for developers and operators, and it's amazingly effective when people find out how fast it is, like "you can literally just hit that button?" Ian warns that some developers get bit by the bug and want to do it all the time.
Ian is @IanColdwater on Twitter, says Twitter is a good place to learn about security, and warns that their LinkedIn says it's only good for phishing. Ian's parting thought: "You don't have to be anybody other than who you are," diversity is strength, and it's important to step into other people's shoes, whether a security person stepping into a developer's or an operator stepping into an attacker's. "Lead with empathy, it's important."
Bridget chats with Ian Coldwater at the devops Minneapolis meetup about their KubeCon North America 2019 keynote.
Video from KubeCon: Hello From the Other Side: Dispatches From a Kubernetes Attacker
Art credit: Sarah Becan - original tweet, threadless store, website
If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf
Bridget talks with all four authors of the forthcoming O'Reilly book Kubernetes Best Practices: Brendan Burns, Eddie Villalba, Dave Strebel and Lachlan Evenson. Brendan writes because "I like to teach," a legacy of once being a professor. Eddie has been at Microsoft for 10 years and wants to spread what the big organizations learn, good and bad, to startups that can't get that help. Dave helps customers succeed with Kubernetes daily and never aspired to write a book, but likes breaking down complex technology. Lachlan wanted to give back to the community, remembering Brendan standing in a hallway in late 2014 or early 2015 answering all of Lachlan's questions, and to write the book Lachlan wished existed in 2015. The cold open is Brendan's line: "It's a powerful tool, but it's also kind of a footgun."
Brendan says the project has moved from something people heard about to something everyone wants to implement, but people struggle with specific tasks, and hands-on help doesn't scale. They've seen "lots of people sort of shoot themselves in their foot." Unlike general introductions to Kubernetes, the book is focused on specific topics, to dip into when working on machine learning or setting up a cluster for a bunch of developers, so it's "a series of short essays rather than a whole put-together book." The 258-page PDF Bridget has in front of them has no narrative flow, which is on purpose.
Lachlan says now is the right time because adoption has grown, the ecosystem has become more complex, and Kubernetes has a sprawling variety of APIs, so the book shows where to start on topics like policy, rolling upgrades, governance and security. Eddie says organizations are already down the path and don't want another step-by-step walkthrough, and the authors tried not to make it a snapshot of one version. Dave says users need to focus on the core concepts and often skip them to over-engineer. Lachlan likes the mix of philosophy, meaning why you'd want policy, and the tactical how.
Lachlan's favorite was Chapter 11, on policy. Lachlan notes "everybody loves hearing Chapter 11 for anything," and says enterprises moving workloads to Kubernetes ask how to make sure workloads conform to policy, whether regulated or just wanting to understand configuration. Bridget notes the chapter covers the open source project Gatekeeper, and Lachlan says it's a Kubernetes-native implementation of OPA, the Open Policy Agent.
Dave's was resource management, which "doesn't sound really exciting at all" but is something users struggle with and affects scaling. The book covers best practices around requests and limits and how workloads behave when capacity runs out. Lachlan says most of the outages Lachlan was paid to handle in the early days came from resource management, as clusters got to 80, 90, 100%, and would have liked to have had the chapter in 2015, to avoid a cluster going into cascading failure at 3:00 AM.
Eddie's was Chapter 9, covering networking, network security and service meshes, which was the most challenging to fit into a concise format. Eddie calls networking the foundation, where little things trip people up, like the move from kube-dns or SkyDNS to CoreDNS, and describes a customer where divisions put Kubernetes in without telling anyone, and then security asked why their controls were gone. Eddie says people want a service mesh as an "easy button" for observability, security and policy, and find it's "Thousands of little buttons that you have to press in the right combination." The chapter describes what all service meshes should do, what to prioritize, and the SMI spec, a common API for those things.
Brendan's favorite, after the first chapter on laying out a service, is the one on developer workflows. Brendan worries that operators love the technology, or it's great for continuous delivery, "but we've made the developers' lives miserable." The chapter covers partitioning a cluster with namespaces, onboarding a new developer, RBAC so people don't step on each other, cluster-level logging and monitoring that's just there, and testing and debugging, since "if it's not easy, people will do less of it, and then you ship buggier software."
Bridget asks whether people in the field actually follow these recommendations. Brendan says people want to, and "the recipes just aren't there, necessarily." Eddie describes a customer creating a new division to build patterns for developers and ops so onboarding is easy and everything is automated. Dave says there's a big cultural impact, since you leave control to Kubernetes that you used to hold tightly, much like adopting a DevOps culture. Lachlan says there's no single right answer, but a set of tools and techniques that have worked, which the book offers as blueprints, so you aren't "left scrounging."
Lachlan's advice: this is a journey and not a destination, there will never be a point where the system is perfect, and people paralyzed by indecision should use best practices to make a decision and keep adjusting. Dave's is to walk before you run, and get good at the core capabilities, network security, resource management and policy before layering on tools like a service mesh installed with one Helm command. Eddie's is to keep an open mind, since practices from on-premises servers and VMs may no longer fit, and the bias of "that's not how we did things" is the biggest challenge in the field.
Brendan's is to understand why you're making every decision, including Kubernetes itself, and not adopt it because "Everybody needs a Kubernetes strategy" appeared in a magazine. Brendan says Kubernetes makes deploying easy without helping you understand the system, and that things hard to manage over time are easy to start, so people assume the ease will continue. Brendan's summary: Kubernetes "is alive," a living thing and not a static one, that "could turn on you at any moment." Bridget calls it the Admiral Ackbar principle: "it's a trap if you think that it's going to be easy." The episode closes with a joke about putting a paper copy in a time machine, or a DeLorean, so Biff can't steal it.
Bridget chats with the authors of Kubernetes Best Practices: Brendan Burns, Eddie Villalba, Dave Strebel, and Lachlan Evenson. (Kindle version available now!)
Dave:
Brendan:
Eddie:
Lachie:
Bridget:
If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf
Bridget records live at devopsdays Philadelphia 2019 with Peter Shannon, a senior software engineer at Instacart and head organizer of the event, Jocelyn Harper, a senior associate software engineer at Capital One who gave the opening keynote, and Tim Gross, a software engineer at HashiCorp working on Nomad, who has keynoted Philly in past years. The conversation is about ethics, data and the responsibility of technologists.
Peter says the program is single-track, in its fourth year, so it has to balance culture talks, technical talks and talks that feel timely. Attendees complain in both directions, that it's too cultural and that it's not cultural enough. Peter tries to make the opening talks lean cultural and relevant to the moment, and if possible frames other talks around that theme.
Jocelyn's keynote was about ethics in technology and the responsibility technologists have. Jocelyn says technologists can seem indifferent to the problems that larger tech companies they use or work for cause for people who aren't technologists. These conversations aren't new, and Jocelyn says social media, Twitter especially, has been a vehicle for people who might not have been heard inside companies.
Tim's earlier talk at Philly turned Conway's Law around: the systems we build and the choices around them also shape our organization's culture and the wider culture. That includes building for reliability so people trust each other, and which technical communities you join, such as a language community with toxic behavior. Tim continued the theme at devopsdays Minneapolis the year before, on whether we build technologies with a good purpose or a neutral one that can be used in bad ways.
Bridget says technologists tend to look at the happy path and ask how realistic it is that something will never be abused. Peter says engineers look at data as inputs to a service, where the inputs can be extremely sensitive and could harm the people the data is about if they got into the wrong hands. Jocelyn says we've reached the point where we take all the data because it comes with it, and should ask "do we really need this data in order to perform the function that we have?"
Tim gives an example from a job a couple of jobs back, a sensor that counted people and detected movement through a space. The designers saw how individual tracking could be abused and designed it so that identifying individuals was impossible, which became a competitive advantage because customers also didn't want their workers tracked. It also detected mannequins moved through a doorway as people, which Peter confirms meant people carrying mannequins in retail. Peter says the general consensus is the opposite, recalling Etsy's "if it moved, graft it" for metrics, and that for customer data it's "Storing bits is cheap." Bridget answers that it's cheap until the lawsuit, and cites a line Bridget can't attribute, that PII is like toxic waste, something you have to carefully corral.
Peter has worked at two organizations with HIPAA data and says the care was amazing, with legal review, access controls and training, but most other data gets a shrug. Tim adds that seemingly meaningless data points can be combined in harmful ways, and says data lakes are a blob of data with models applied that aren't magic, since "we're really encoding like, all the biases of the people who are doing that work into those models." Tim's example is parole models, and "oh, weird, it turns out that that algorithm is completely racist." Peter adds "And they blame the algorithm, not the people who wrote it." Jocelyn says the question is also why the data is collected, and that "just collecting data just to have it" isn't a good enough reason.
Tim says some technologies are neutral, like Kubernetes, and all can be turned to bad ends, but asks whether there are any good uses for facial recognition. Jocelyn says many engineers see it as a job and don't think about the impact, and hopes people take away that "it's your responsibility to think about that," and that not being on the executive board doesn't mean you can't have an opinion about what your work is used for. Bridget says you don't need a dramatic splash, and that if you think about operability and how systems break, that extends to data breaches and privacy, such as a provider who emailed that everything had been in an S3 bucket open to the world.
Peter says it comes down to money: if a breach would be company-ending it gets bolted down, but if the cost of fixing is more than the cost of the damage, it doesn't. Tim says companies have been allowed to externalize these costs, like pollution, and that if technologists don't fix it themselves, lawmakers may eventually decide how technology works, which probably won't be a good technical solution. Jocelyn thinks it will take federal law with real penalties. Bridget notes GDPR has penalties with teeth, and that some US companies respond by dropping Europe. Tim says California enacted a similar law, and most US companies won't cut off California. Peter says a toothless law won't work, and fines need to scale per infraction, not be a one-time fee.
Tim says technologists have a lot of privilege and power, and "if I were to ask who here is hiring, everybody would raise their hand." That power includes choosing who to work with, which projects and who the customers are, and Tim says it's time to exercise it together. Jocelyn says to start at work by being the squeaky wheel, and outside work, to start on Twitter, where Jocelyn began before podcasts and talks, and says "Twitter activism does work." Peter echoes the "vote with your wallet" idea for where people work, acknowledges leaving a job on principle is hard for some, and says that if people banded together and left companies that don't meet their standards, it "would really, really send ripples through the industry."
Bridget chats with guests Peter Shannon, Jocelyn Harper, and Tim Gross in front of a live studio audience at devopsdays Philadelphia 2019.
If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf
From the publisher's feed