Arrested DevOps

Arrested DevOps

By Matt Stratton, Trevor Hess, Jessica Kerr, and Bridget KromhoutTechnologyTech News
Download on the App Store

Arrested DevOps episodes

  • starting a new devops job

    Tonight's episode featured a special guest host, Julian Dunn of Chef.

    Our panel:

    Ryn Daniels - @rynchantress - a web operations engineer from Etsy.
    Jake Champlin - @grubernaut - an operations engineer from Minted. (Younger than Trevor!)

    Matt asks Ryn what it was like going to a place like Etsy - the gold standard! - what they expected, and what it was like when they got there. Ryn says that years of reading codeascraft and seeing etsy engineers speak at conferences meant that they were excited to start but also felt a little impostory. "Everyone knows that devops is the internet and the internet is cats - I show up my first day and my desk is covered in cat pictures."

    Jake talks about how devopsdaysPGH was his first tech conference, where he got connected to the community and met his future boss Alex Nobert. Julian asks how Jake got connected to Minted, and Jake talks about Bridget introducing him to Alex and how Minted seemed to have a great culture. Trevor asks what stands out about Minted's culture, and Jake mentions that it's woman-owned and has a strong focus on the product, and how everyone's motivated to make a quality product and it's reflected in your infrastructure.

    Bridget suggests this might be the inverse of Conway's Law, while Matt asserts that most "inverse Conway's Law" discussions just end up proving it. Jake mentions how the community of Minted votes on some products, and the engineering team tries to bring that same excellence throughout.

    Bridget asks Ryn how Etsy's engineering practices influence or are informed by the culture of the company. Ryn mentions Etsy's B-Corporation status, which means they are focused on social good, transparency, giving back to the environment - and that's reflected in a lot of what they do, from codeascraft to blog posts to involving sellers in the community.

    Both guests give a short version of what their company does; Ryn says that aside from the monitoring and sending everyone to Velocity, Etsy is the largest marketplace for handmade and vintage goods. When Jake explains Minted, it's revealed that as "lazy podcasters", we didn't realize that Minted was as artist-focused as it is, providing artists a platform and community - Bridget thought it was primarily for paper goods, while Matt thought it had something to do with money.

    Bridget asks Ryn about their decision-making process to decide that Etsy was right for them. They said that even though Etsy folks they met at Velocity didn't know them, they were welcoming and never condescended to them. (Also, pink-haired thought leadership.) They also admired how Ian Malpass was mentioning at devopsdaysMSP how he wanted to put together a class for effective male allies inside Etsy.

    Similarly, when Jake went to devopsdaysPGH, the way people were open and welcoming is what made him know he was in the right place. (A restaurant-related digression led to a shout-out to Jon Cowie's knife-spork.)

    Matt mentions having the feels and being excited about getting a chance to work with Julian Dunn! and how when starting a new job, when you feel intimidated, people in a good culture are going to reach out to you and make you feel welcome. Julian asks our guests what the experience of starting at one of these devops gigs was like.

    Ryn talks about how starting at Etsy is less-stressful than smaller jobs in the past (since payroll isn't an open question) and then mentioned the engineering rotations that allow new people to "bootcamp" with other teams, to increase understanding across the site. They also mentioned the "first push" program, where everyone (even non-engineers) learns to deploy a change to the site.

    Jake's working remote, and his onboarding process was like being thrown into the deep end of the pool. He felt impostor syndrome at that point. Bridget agrees that it can be an intimidating feeling, and asks Jake how he deals with that. He says that he's working with really smart people and is happy to learn new things every day.

    Trevor asks about cultural cues. Ryn says that they have a chat-heavy culture, and because there are so many remotes, even the in-person team will communicate as if they are remote, so you get to know what's going on with the team and what their favorite cat gifs are. Trevor mentions introducing hipchat to a client and how it caught on quickly. Jake mentions that although only ops and qa are remote, everyone at Minted uses hipchat. Matt believes it's frustrating when there are people in an org aren't on chat, and Jake agrees that having the whole org there (HR/payroll, etc) makes communicating easy, even as a remote employee.

    Ryn: "I do get to sit ten feet away from John Allspaw, so I've got that going for me."

    Matt: "If you're playing the Arrested DevOps drinking game, that's a drink."

    Julian asks what are some of the downsides of a chat-heavy culture. Bridget mentions that at DramaFever, people have sometimes found chat to be distracting, if people are wanting your attention all day, and Ryn points out that at Etsy, it's considered okay if someone needs to be heads-down working on something, either turning off notifications or signing out of chat. Trevor mentions that working with a global team can mean unintended awakenings from 3am @-mentions, while Ryn and Jake are both in cultures that encourage not having chat on their phones. Bridget says that at DramaFever, the solution is spaces in someone's name so as to not alert them during off hours.

    Jake talks about how being in an oncall rotation instead of being oncall 24/7/365 is great, and after a week of oncall, at Minted they get the next Friday off (and their co-workers will kick them out of chat if they join). Bridget asks Ryn about their blog post on work-life balance, and Ryn says they wrote that about disconnecting, and Etsy also has a culture of asking people to take care of themselves (where they threatened to take away their VPN access because they were working while sick).

    Julian, who reveals he is stranded in Chicago and that's why he is on the show, turns the conversation back to finding good roles and asking what role a recruiter plays. Ryn says that an Etsy person being oncall and troubleshooting during a Sysdrink meetup in New York attracted a number of applicants. That, codeascraft, speaking at conferences - showing what an org does, instead of just saying "we're hiring" - works better.

    Bridget asks them what makes a job a place they know is right for them. Jake says the culture, and working on interesting things with smart people. Ryn agrees, and says that their first week at Etsy they got to start contributing right away. Bridget points out that new people have a power that people who have been there longer can never get back - the power of not knowing how things are "supposed" to be (and the ability to write documentation to better serve a "don't know the answer" POV). Ryn says that breaking Nagios Herald their first week (and then they also got to experience Etsy's culture around blamelessness.) Jake is in a more greenfield situation, and so he was able to start contributing immediately out of necessity.

    Bridget asks for the panel's best advice for people who are interested in a "devops" job, want to find such a job, etc. Apparently the answer is having Bridget introduce you to people and/or convince you that you're totally good enough to work somewhere. [Note from Bridget: this may not scale.] Both Jake and Ryn point out that connecting with the community on Twitter is a really valuable place to start making connections. Matt points out that liking your co-workers isn't so much "nepotism" - you don't have to party with them - but you spend a lot of time with your co-workers, so you're going to want them to be people you don't hate. Matt says, "It's not know the 'right' people, it's just 'know people'." Jake points out that at PGH, Mark Imbriaco and Ben Rockwood spoke to him as if he was a friend, even though they are prominent in the community. Ryn encourages everyone to use Twitter.

    Checkouts
    Ryn
    • git troubleshooting
    • tmate - terminal sharing
    • Jake
      • Running Linux on a Mac
      • Helena Nelson-Smith on Mental Overload
      • Julian
        • http://usersknow.blogspot.com/2015/02/your-job-is-not-to-write-code.html
        • http://www.isaacchansky.me/days-since-last-new-js-framework/
        • Trevor
          • Laphroaig PX Cask Triple Matured Scotch
          • The Book of Mormon
          • How we upgrade a live data center - http://blog.serverfault.com/2015/03/05/how-we-upgrade-a-live-data-center/
          • Bridget
            • Fat Bike Birkie - http://www.birkie.com/bike/events/fat-bike-birkie/
            • Bike camping in Napa
            • 30 Days of Biking: http://30daysofbiking.com
            • Matt
              • The Goat Farm - Enterprise DevOps podcast by Michael Ducy and Ross Clanton.
              • DevOps Checklist - from our buddy Steve Peirerra
              • 56 min
              • Docker! Docker! Docker!

                James Turnbull describes Docker as “a solution that is built by people to be usable by people, as opposed to some of the previous containerized solutions which were built by engineers to be usable by a very small subset of other engineers”.

                When you might want to fly less… “when you start to recognize the airport lounge staff and the flight attendants on the New York-San Francisco route and they start to recognize you.”

                James wrote The Docker Book to allow people of varying skill levels to quickly understand how to use Docker and what the practical applications could be for them. It’s intended to be a practical how-to guide.

                At Kickstarter, developers have a dockerized replica of production on their laptops.

                Matt asks if Docker can be only used in completely new deployments designed for Docker from the ground up. James points out that if you have existing infrastructure tools, it’s simple to create Dockerfiles from them.

                The night before 1.0 launched at the first Docker conference in mid-2014, James removed all references to “don’t use this in production” from docker.com.

                James mentions that Fig (soon to be renamed to Compose) helps with modeling multi-tier architectures locally.

                James says, “People kinda forget the past and go, “oh my god Docker’s a pain in the ass to use”, and I’m like “compared to what, exactly? Compared to your previous build, or compared to you shipping around 10 ISO files and running Vagrant and 20 VMs on your local machine?”

                He continues, “It wasn’t that long ago that the dark ages were real. I’m not suggesting that Docker’s a panacea, but it’s certainly a step in the right direction.”

                James points out that something like Elasticsearch does well in Docker, since “it’s a bit of a fiddly thing to build, with the right version of the JVM, right version of Elasticsearch, prepping all the data, etc”.

                James highlights continuous integration as a “sensational combination” with Docker.

                On the controversy, James points out there will always be hype and people claiming “this is a revolutionary technology that will cure world hunger”. He says, “I’m fond of saying that Docker is a powerful tool to help you in your development life cycle [...] not every workload in your data center is well-suited to Docker.” James doesn’t make technical architecture decisions based on the writing of tech journalists or blog posts, but rather by testing and evaluating the relative merits of a given solution.

                In the case of Graphite, James would run carbon-relay and carbon-cache inside Docker containers, but he’d point them at a physical machine with SSDs to actually write the whisper files.

                Matt read a blog post and reactions on reddit and wanted to see what James thought of the concerns around security and operability. James points out that empathy for developers is something sysadmins need to cultivate, because you don’t manage infrastructure for infrastructure’s sake.

                James points out that the main reason developers ship code that doesn’t work in production is that they have no fucking idea what production looks like because there’s this grumpy asshole that manages production and they’re terrified to go ask them a question. Bridget says that as such a former grumpy asshole, she’s much happier when the devs aren’t afraid to talk to her.

                James mentions that Docker containers are not virtual machines and should not be used to separate security concerns, and you should secure the host the containers are running on.

                Matt: “I’m not suggesting that this [security concerns] is why DOCKER BAD…” Bridget: “Race conditions with devicemapper is why DOCKER BAD.”

                James: “[PCI/DSS] is a low bar. If you followed simply the regulations for the compliance stuff that related to PCI/DSS, you would be running a massively insecure system.”

                James points out that “owning” the standard gives one access to the marketing around an ecosystem. He also thinks that even if Rocket is a better technical solution, Docker has more traction.

                Bridget: “So when I feel ranty about Docker and devicemapper, I should submit some pull requests.” James: “You should talk to Michael Crosby... Michael Crosby is currently in San Francisco somewhere going you motherfucker.”

                James sees Amazon and Microsoft’s embracing of Docker as a great driver of revenue towards these cloud providers, if it gets developer code to production faster. They aren’t following hype; there are transparently obvious business reasons to do it.

                In terms of skating to where the puck is going to be, James suggests looking at orchestration, software-defined networking, software-defined data centers - people building that sort of thing with Docker components. Docker Compose, Docker Swarm, people moving up the stack to manage different levels of abstraction.

                James: “I challenge you to find a LAMP stack site where 80-90% of the configuration files aren’t identical - our secret knowledge of what to tweak isn’t as valuable as we think it is.”

                Check outs
                James
                • Jason Dixon’s Monitoring with Graphite book
                • The Art of Monitoring - James's upcoming book
                • Bridget
                  • Spent a week with my Philly & NY co-workers, went to the 3rd annual DramaFever Awards show, sang more off-key karaoke.
                  • Docker to Ducy Chrome plugin
                  • Matt
                    • Was in PDX for the first time last week for Agile Open Northwest. Led an improv open space. Got a tattoo. Met the founder of Voodoo donuts winding a clock. Such hipster. Very Portland.
                    • kitty gem
                    • Better Call Saul - new AMC show
                    • 1 hr
                    • DevOps in a Microsoft World with Jessica DeVita and Jeffrey Snover
                      The GUI as Strength and Weakness

                      Jeffrey Snover, a Distinguished Engineer and lead architect for Windows Server and System Center at Microsoft, found DevOps through John Willis's podcast and Gene Kim's book and "fell in love with it." Jeffrey says a lot of it was familiar, since Jeffrey wrote the Monad Manifesto behind PowerShell in 2002, and that it echoes the quality revolution of the 1980s at Storage Technology. Jessica DeVita, a technical evangelist at Microsoft, runs IT camps where Jessica talks to traditional enterprise IT people about configuration management and version control. Jessica says the IT pros at those camps are frustrated after being promised solutions that haven't worked, can't tell "the wheat from the chaff" among tools, and often work in environments that aren't supportive of the person at the console.

                      Jeffrey says Microsoft's strength in GUI tools is also its weakness, because "it's very hard to share a bunch of mouse clicks." Deployment guides run to huge books of screenshots that say click here, and the PowerShell equivalent is a page and a half. Some people love the change, and others hoped they were done learning. Jeffrey's advice: "If you don't want to learn anything new, get into the lumber business," since in IT "you're riding the tiger."

                      The Body Follows the Head

                      Asked about Microsoft's turn toward open source, Jeffrey says "organizations are like gymnastics: the body follows the head." A new CEO brought a fresh approach of meeting customers where they are, and for people who'd been trying to do this for years, it had not been a friendly environment before and now is. Jeffrey's way of explaining it to skeptics is Azure, where Microsoft makes more money from 10 Linux instances than from 2 Windows ones, and its services are REST APIs usable from anything. Microsoft, Jeffrey says, is "becoming a no-adjective software company." Jessica, who joined a couple of months after the new CEO, says the open source community has cultural lessons to teach the enterprise.

                      Jeffrey says this is the beginning of a long journey, with things already open like OMI and Desired State Configuration, and that the only reticence Jeffrey has seen is "here's all the things on my plate." Jeffrey adds that Microsoft is "incapable of sustained error": it screws up constantly, but people call each other out without firing anyone, which works like an immune system. Matty ties it to the blamelessness episode, number 28.

                      Why Not Group Policy or SCCM?

                      Matty says every way Microsoft has told Matty to configure Windows servers has been tried. Jeffrey agrees that Group Policy and SCCM exist and are good for enterprise client management but not for the data center. Jeffrey also recalls a Passport server admin who set a policy with no way to tell which servers got it. Desired State Configuration is focused on a DevOps way of configuring servers, and says the architecture won't be compromised to serve clients. An executive once told Jeffrey to solve server configuration management, which Jeffrey called impossible, since every group thought of itself as its own CTO. Jeffrey's guide is the DevOps line that "scale times complexity exceeds our skill set," so it has to be "simple, simple, simple."

                      Jeffrey wanted a platform that could configure everything in the data center, including Linux, routers and storage, and that others could build on. Why not Chef or Puppet? The core difference, Jeffrey says, is that Unix is document-oriented and Windows is API-oriented: bringing awk, grep and sed to Windows, which Jeffrey once tried, didn't help, since they don't work against the registry, WMI or Active Directory. The exception is IIS, which is why Matty says the IIS cookbook is so good. Matty adds that ConfigMgr is "a big database," which you can't version and can't treat as code, and that a common language between devs and ops lets people pick up 80% of Chef or DSC quickly.

                      Getting People Off Click-Next

                      Jessica says the challenge is Windows admins who've lived in the GUI and never developed command line skills, a generalization. Jessica likes that Server 2012 R2 shows the PowerShell before the Finish button, so you can copy it and learn, and says more of that is needed. Trevor asks about places still on 2003 and 2008, and Jessica says "you should just already love PowerShell," calling it a different podcast. Matty's line is that it's called PowerShell, not PowerScript. It's not a scripting language but how you interact with a system, so start by telling someone to "crack open a PowerShell prompt and type this in."

                      Jeffrey points to Don Jones's PowerShell in a Month of Lunches and a Microsoft Virtual Academy course, and tells managers to promote and reward the people who are moving you toward repeatable, automatable IT: "If you're going to retire in the next 3 years, like, forget it, you know, just click next and learn, you know, practice Bridge." Matty says there are still jobs for AS/400 admins, but it would drive Matty crazy. Jessica says automation buys time, "the real currency," and captures what a brilliant sysadmin figured out so nobody has to keep rediscovering it, "a way to really version control your culture," which Jeffrey answers with "I like that."

                      Undifferentiated IT and Healthy Fear

                      Jeffrey doubts the click-next jobs will last, since if you're offering undifferentiated IT, the cloud will offer it cheaper, more securely and with better data protection. But if you understand the mission and provide differentiated IT, "you're printing money for your company," and they won't take a generic cheeseburger from the cloud. Jessica says the fear should be of more interesting things: automation is what people should learn, and companies will hire you to automate their infrastructure.

                      Jeffrey goes further: "fear, uncertainty, and doubt, these are your friends." Jeffrey tells of a karate student whose hands kept dropping until the instructor decked the student once. Jeffrey is a college dropout who comes in every day with imposter complex, and responds by working hard and performing. Jessica dropped out too, and Trevor says dropping out seemed to mean never succeeding in technology. Matty: "Is there ever anyone on this show that has a degree?"

                      The Best Tool, or the Safe Bet

                      Matty asks how to help people who avoid non-Microsoft tools, like Lync versus HipChat for ChatOps or SCOM versus Nagios, on the "nobody got fired for buying IBM" theory. Jessica says chat culture matters more than the tool, and uses Yammer. Jeffrey says some parts of Microsoft got DevOps in focus earlier than others, that as more teams run their own services they demand better tools, and that with about $10 billion a year in R&D, "we can move fast." Jeffrey teases what's coming.

                      Jeffrey also argues that the best tool has to be around next year. Jeffrey worked at Digital Equipment, Apollo, Greystone, Royce Data Systems and Storage Technology, and "zero of those guys are around." Jessica says weighing whether a tool would last was always part of the job as a consultant, checking who the founder is and where it's hosted. Matty's answer is to pilot: make a small experiment with people who are on fire about the goal, the way Agile was introduced, and Matty cites the GE story from ChefConf. Jeffrey offers a dissent about letting teams choose freely, because a developer's Erlang, or an inherited product at Digital written in 18 languages including Ada ("Brad took the night course in Ada"), is a mess when that person leaves. Jessica says a language is a different level of impact than a chat tool.

                      The WinRM Question

                      Matty asks on behalf of Brian Barry of the Food Fight Show why you can't copy a file to a server over WinRM. Jeffrey says someone is prototyping it, and that the performance can be terrible compared with SMB, and tells Matty to come talk at Build and Ignite. Matty says it will do as a minimum viable product.

                      Discussion Outline
                      What are some of the challenges traditional Microsoft IT Pro’s deal with moving to a more automated DevOps pattern?
                      • Jessica:
                        • Hard to tell which tools are really going to make their lives easier.
                        • Are the cultures of the companies benefiting the human side of the IT Pro?
                        • Jeffrey:
                          • Because Microsoft has great GUI tools, they become the biggest strength and weakness of the DevOps/IT-Pro
                          • The process of using a GUI is much harder to replicate in documentation. Because most of the community uses powershell commands, Microsoft IT-Pros really need to get on board.
                          • IT-Pros are never done with learning. If you don’t want to learn anything new, get into the lumber business.
                          • Microsoft has been making more open source integration moves, and changing philosophies to accept the Open Source community. “What up with that?”
                            • Jeffrey: “The body follows the head”
                            • It helps when you have a leader with a fresh approach who focuses on customer service and helping users within the community
                            • Jessica: It is really exciting to get behind a leader that is welcoming to the communities.
                            • It is refreshing to see Microsoft becoming a software company, not a “Windows software company”
                            • Microsoft wants you to be successful. Tools such as RESTful APIs are becoming available across all OSs.
                            • What is the acceptance level of the OpenSource movement within Microsoft?
                              • Jessica: Whatever you’re running, we can host it for you
                              • Traditional Configuration Management in Microsoft has been difficult. What are the plans?
                                • Steve Morowski (http://stevenmurawski.com/) has good info for those interested in DevOpsing with Windows.
                                • Current Microsoft tools are really good for enterprise, client management. Not so good for data center management.
                                • We need something different, that is simple, and usable.
                                • The problem is, everyone wants to do configuration their way. They want to be the CTO of their servers.
                                • Jeffrey describes the creation, and idea conception of a Microsoft Configuration Management platform that takes into account the deep differences between Linux, Unix, Windows. Describing different tools currently available, their faults, and how they might be able to connect them for modern, DevOps oriented, Configuration Management.
                                • The ability of chef and puppet, etc. are beneficial because of the ability of devs to pick it up, version it, and insert small parts of just what they need into the configuration.
                                • Jessica: We are getting to the point where Microsoft DevOps engineers are adapting the powershell. Until the powershell is adopted by IT-pros, modern DevOps tools will be a difficult push.
                                • You should already love powershell.
                                • How can people get more comfortable with powershell?
                                  • Matt: It is not a scripting language. It is the way you interact with a system. Don’t write scripts in bash, write commands in bash that emulate the scripts.
                                  • Jeffrey: Don Jones: Powershell in a Month of Lunches (http://morelunches.com/2011/04/01/learn-windows-powershell-in-a-month-of-lunches-1st-ed/) Step by step people get it, or they don’t. Managers really need to promote and reward the people giving you the IT that you want.
                                  • Poweshell makes your environment repeatable, automatable, stable, etc. It is the future of the IT pro, and people must adopt it.
                                  • Are we automating ourselves out of jobs?
                                    • Jeff: The cloud is a great, cheap place to offer undifferentiated IT, however, if you can provide differentiated IT you are practically printing money vs. the cloud.
                                    • Jessica: We do need a healthy fear. Not of automation though. Be scared of more interesting things. You need to learn automation.
                                    • How do we work with Microsoft when its just not the best for DevOps-ing?
                                      • As more people us Microsoft, the more Microsoft changes. Jeff discusses the many ways in which Microsoft is using flexible R&D to make a push for DevOps tooling, as well as some tools coming down the pipeline.
                                      • Jessica: When choosing a tool, the longevity of the tool and the community around it is critical.
                                      • Why can’t I copy a file to a server using WinRM?
                                        • Jeff: Come talk to me at ‘Build and Ignite’.
                                        • Checkouts
                                          Jessica
                                          • The SoCal Linux Expo - Scale13 - a DevOps day Feb 20th - she has a discount code
                                          • The Field Guide to Understanding Human Error (Dekker)
                                          • Lean Enterprise book (Jez Humble)
                                          • Jeffrey
                                            • Hardcore History podcast -  I’m in love with this podcast. Dan Carlin is an awesome storyteller.
                                            • Brain Science Podcast - (Cool podcast about the brain. I was just telling someone about this today)
                                            • http://www.microsoftvirtualacademy.com/training-courses/getting-started-with-powershell-3-0-jump-start This is the start of a 2 day training session on using PowerShell. It is one of the most widely viewed jumpstarts ever.
                                            • http://www.leeholmes.com/blog/2011/04/01/powershell-and-html5/ One of my all time favorite PowerShell scripts.
                                            • Trevor
                                              • FCC Ruling on broadband
                                              • Windows 10 on Raspberry Pi 2
                                              • Matt
                                                • Kitchen-windows is almost a thing! If you want to play with it, check out the Windows cookbook at http://github.com/opscode-cookbooks/windows
                                                • BitTorrent Sync - http://www.getsync.com/
                                                • Yes, Please by Amy Poehler
                                                • 1 hr 20 min
                                                • Hiring in a Post-DevOps World
                                                  A Job Description Nobody Agrees On

                                                  The episode was prompted by Josh Hertz, a self-described professional rabble-rouser and the sysadmin for a podcast, ranting to Matty about job postings that want ninjas and rock stars. The panel is Mike Fiedler, Director of Technical Operations at Datadog, and Jill Jubinski, who manages technical recruiting at DigitalOcean. Mike says the biggest hiring problem is that the job description is "entirely too wide and hazy," because everyone uses words nobody agrees on: we want 20 years of Docker experience and every DevOps tool, and someone who's capable of everything and amazing at anything. Post for DevOps when you want a build engineer, and you get contention. Jill says the word means support at some companies, engineers at others, and developers at others, so you have to understand what the job actually needs "and not just have all those crazy buzzwords."

                                                  Matty repeats the observation that nobody in Chicago hires sysadmins anymore, only DevOps engineers whose descriptions read as sysadmin jobs. Matty once spent 25 minutes in a cab explaining to a contract recruiter why the posting was bad, and says the name doesn't matter, just describe it: "no, we use Jenkins."

                                                  The Sysadmin Badge

                                                  Mike says at some point in the last decade the title system administrator "became a dirty, dirty badge." You can't study for it, Mike adds, and the sysadmins of the corporate mail and web servers "hated everybody because running a mail server is usually not very much fun." When companies needed operational mindsets, they reached for a new term. "Not only can we be a DevOps, we can be a DevOps ninja."

                                                  Josh says the start was as a sysadmin, moving to systems engineer when designing systems to scale began, and that DevOps suggests you're doing continuous integration and configuration management. Matty says system engineer wasn't just a nicer word for sysadmin, but that sysadmin became a "server janitor" title.

                                                  Ninjas, Rock Stars, and Heroes

                                                  Matty cites a Reddit user on r/devops: "ninja and rock star are code words for you will be one person doing the job of an 8-person team." Mike says a rock star is one person in the spotlight supported by a band, so what Mike wants is "a tenor in a choir." Josh suggests a session musician, and Matty says you don't want "a DevOps Axl Rose." Jill says the words conjure hiring that one difficult person on your team, and that what teams need is a team player.

                                                  Jill and Mike discuss hiring someone a step above the team. Jill likes it if they share knowledge and raise everyone, and Mike's point is that it's how you frame it to the team: if you're telling everyone else they're now second string, that's a terrible message. Matty says a 10th Magnitude ad Matty came across was rewritten to say they're not hiring ninjas, rock stars or unicorns, just really good developers, winking at the cliché.

                                                  Matty adds that rock stars signal hero culture, and points to Jennifer Davis's talk "From Hero to Zero." Mike says there's a time for it: a three-to-eight-person startup's first ops hire may need the hero who'll work 120 hours a week. At a company with a 30-person operations team, though, the hero is "quickly going to be not as popular amongst their peers."

                                                  Recruiter Fail

                                                  Jill reaches out to 15 to 20 engineers in a given week, each with an email special to them, and says mass emails come from agency recruiters, where "it's all about numbers." Josh asks if targeting works and Jill says absolutely: people are more receptive, and the aim is to build a network so that when your friend is looking, "you'll think of a non-shitty recruiter."

                                                  Matty's LinkedIn says Matty loves the job and has zero interest in any opportunity, and a recruiter who had clearly read the profile replied that people write that to lay low for their employer. Mike says recruiters who do a cursory Google and then send a Ruby on Rails front-end job to someone who isn't one are basically sending spam, and that nine times out of ten those emails get ignored. Jill says engineers on the team forward funny recruiter emails, and that Jill sometimes cold-emails coworkers a Drupal job for a laugh. Matty stays in touch with the polite recruiters who ask for five minutes of help, even if they've never placed Matty.

                                                  What Candidates Get Wrong

                                                  Jill says resumes full of buzzwords and things you touched but didn't dig into get figured out, and Mike says to be honest: three lines of Perl doesn't make it a skill, and if the interviewer asks about something esoteric and you flounder, the phone screen is "going to take a nosedive." Trevor adds that if you say you don't know something in the first interview, don't act like an expert an hour later, "because the interviewers do talk to each other." Josh interviewed someone who listed expert on the OSI model and named three of the seven layers, and Matty's response is that in baseball that would be a good batting average.

                                                  For interviews, Matty tests whether you understand things, not whether you remember every tar flag. Trevor learned that whiteboarding isn't about the right answer but how you get anywhere, and Mike's question is one with no right answer where Mike keeps throwing monkey wrenches to see how you handle adversity. Jill says smart organizations hire for critical thinking and problem solving, since tools like Chef and Go change in three months. Josh says DevOps is knowing what you don't know. Jill's pro tip: "If you don't know something, admit you don't know it. Don't make it up."

                                                  Culture Fit and the Resume

                                                  Matty asks how to convey a DevOps mindset in a resume. Josh says listing "problem solver" and "self-starter" doesn't help, and Mike says it comes out in interviews, where the interviewer poses the right questions and the whole team talks to the candidate. Trevor points out that a bad resume never gets you in the door, and Matty's answer is to skip the adjectives and show what you did, such as any experience working across silos.

                                                  Jill wants to see open source contribution, and Matty says you can at least do something small like whitespace fixes. Mike says junior DevOps people rarely have the experience, so Mike would read an objective statement as why you want ops and not development. Mike suggests tailoring several versions of your resume to the company, and Jill's related pet peeve is a candidate whose materials say they want to work at Google when Jill works somewhere else. Josh admits Josh's own summary reads "a DevOps rock star, 10x pirate ninja," which is "100% snark."

                                                  Check-Outs
                                                  Jill
                                                  • Beer: http://www.beeradvocate.com/beer/profile/45/680/
                                                  • Newest obsession: Rowing class: http://www.cityrow.com/ (there are probably many others outside of NYC)
                                                  • Josh
                                                    • I’m officially the SysAdmin for Caustic Soda, which is a podcast about horrible things. It’s like Mythbuster for all that is gross and horrible. Now that I think about it, I can’t recommend it.
                                                    • Influxdb is an open-source, distributed, time series database with no external dependencies. Currently in Alpha (v0.8.8) it's going through a major refactor, the "production ready" version v0.9.0 is due out this month. The ease of installation and setup is what sold me on it. I think it has a lot of potential, so we'll see how v0.9.0 looks when it's released.
                                                    • The Bruery:  Briefly mentioned. Small batch, high quality beers out of Placentia, Ca. Their barrel aged and sour beers are outstanding. I recommend Loakal Red as the gateway beer for those that like IPAs.
                                                    • Mike
                                                      • http://www.opsschool.org
                                                      • Tubes: A Journey to the Center of the Internet, Andrew Blum
                                                      • The No Asshole Rule: Building a Civilized Workplace and Surviving One That Isn't
                                                      • Matt
                                                        • Elevate - brain training app
                                                        • Doing the fitbit thing again. We can be buddies. http://mattstratton.com/fitbit
                                                        • Trevor
                                                          • The Decemberists: What a Beautiful World, What a Terrible World. They're back!
                                                          • Microsoft reviews pull requests for C# from github
                                                          • Walking around your neighborhood. Discovered a new spice shop and comic shop by wandering North.
                                                          • 0 min
                                                          • Incidents and Accidents: Examining Failure Without Blame

                                                            Dave is at Next Big Sound, which does analytics for creative industries, and he’s seen a few orgs handle failure well, and a lot of organizations handle it poorly. He got interested in blameless postmortems and human factors in discussions with John Allspaw of Etsy, and Allspaw influenced him to read the work of David Wood and Sidney Dekker on human factors. He is writing a book for O’Reilly called Being Blameless.

                                                            MCR works at Etsy now, but has spent a lot of time consulting at various firms where he’s seen failure handled with blame. He points out what Rt. Lieutenant Colonel Scott Snook said in Friendly Fire, a book about when two US helicopters were accidentally shot down, that failure is part of complex systems.

                                                            MCR: “I work at Etsy, and that’s what we do - we examine failure as a learning opportunity.”

                                                            Dave is running his next workshop on Awesome Postmortems in NYC on February 12th, in which

                                                            Dave: “Sidney Dekker’s Field Guide to Understanding Human Error is probably the most important book for people like us, meaning people that are in the IT world - it’s very accessible and gives lots of examples from fields outside of IT, but they’ve very relevant to what we do.”

                                                            MCR: “Failure is gonna happen. It’s not a matter of if something is going to fail, it’s a matter of when it is going to fail.”

                                                            MCR mentions the different categories of failures - those that “fail closed”, that are easy to detect, like disk filling up, and “fail open” - the surprises. He mentions some of the techniques Etsy uses - an IRC warroom, Vidyo video chatting, to resolve an immediate issue. After the immediate issue is solved, the learning begins.

                                                            MCR: “We celebrate failure as much as we celebrate success here. [...] The three-armed sweater is given to the person who most spectacularly impacted the website in the year.”

                                                            On the topic of why to do a blameless postmortem, MCR points out that it’s for learning, and there are both technical and human factors. Dave points out that blaming a person short-circuits the learning. Claiming that a person is the cause of the outage feels like a good story, but it’s not true.

                                                            Dave discusses root cause and mentions Allspaw’s excellent blog and a specific post about there being no such thing as a root cause, and Dave disagrees. He believes that outages are caused by change, and the systems with which we work are fundamentally changeable. “The impermanence of systems is the reason that they both function and malfunction.” Mike counters by saying, “Is there really a root cause for something that failed? If a hard drive dies, it’s the same hard drive. It hasn’t changed.” They both agree that it’s a philosophical rabbit hole.

                                                            MCR notes that as Etsy grows, they’ve found that user-impacting, service-degrading issues are when they do postmortems, and even if not user-impacting, if they can learn from a failure it’s worth doing one. Dave says, “The more we learn about the complex systems within which we work, the better we’re able to operate them.”

                                                            Within a week or two, according to Dave, is common practice of a time in which do the postmortem. MCR mentions that it’s important to write down the timeline almost immediately, definitely within a day or two, but doing it while someone’s amygdala is still triggered (and they are upset) is too soon. Dave points out that the facilitator of a postmortem sets the tone, including reminding people of hindsight bias, and at Next Big Sound they use a specific framework document which Dave will share. He also mentions defusing stress with empathy and humor.

                                                            On the topic of evaluating anything you do, MCR mentions that Etsy created Morgue because any department across Etsy can apply these techniques to learn. Dave points out they do retrospectives as well as prospective review at Next Big Sound. MCR says Etsy does both an architectural review and an operability review ahead of time. Dave mentions that answers in prospective reviews can be biased in a positive way, whereas in a “premortem” we imagine things going badly, and try to determine what could lead to that: in essence, harnessing hindsight bias to work for us.

                                                            Bridget forgets what decade it is and claims to have seen a presentation at devopsdays 2003. That would have been a nifty trick, since the first one was in 2009. :)

                                                            Check Outs
                                                            Dave:
                                                            • Field Guide to Understanding Human Error, Sidney Dekker - new edition just came out Dec 28th, 2014!
                                                            • Mike:
                                                              • http://www.bsr.org/en/our-network/member-list
                                                              • http://solutions.3m.com/wps/portal/3M/en_US/NA-DataCenters/DataCenters/Solutions/EfficiencySustainability/ImmersionCooling/
                                                              • Etsy's postmortem tool Morgue
                                                              • Bridget:
                                                                • Did a lot of reading last week! Beside the aforementioned Dekker book, I read Web Operations (edited by John Allspaw and Jesse Robbins), Jon Cowie’s Customizing Chef, and the first three chapters of Jason Dixon’s Graphite book. All good stuff.
                                                                • The Boundary Waters Canoe Area is great in the winter too: I like Gunflint Lodge: http://www.gunflint.com/
                                                                • Trevor:
                                                                  • I was on vacation and delightfully disconnected. It’s been pretty awesome. Got a new Kindle and have been reading Game of Thrones before Matt accidentally (though at this point it’s my fault) spoils something.
                                                                  • Set up kegbot at our new office, will be doing it’s grand opening later today :) Metrics about office beer / root beer consumption to come!
                                                                  • Matt:
                                                                    • Been on vacation, which is great. Doing the Dadops thing although I was sick for most of it, which was not delightful. However, I did see Big Hero 6, and Trevor was right about that.
                                                                    • Book I’ve been listening to on audiobook is The Challenger Sale by Matthew Dixon and Brent Adamson http://www.amazon.com/The-Challenger-Sale-Customer-Conversation/dp/1591844355
                                                                    • 0 min
                                                                    • A Year of ADO
                                                                      How It Started

                                                                      For the year-end episode there are no guests, and Bridget Kromhout, the newest host, asks Matty and Trevor to start from the top. Matty says the idea started as a DevOps 101 blog, a friend came up with the name Arrested DevOps, and podcasts had a gap for people just starting to learn about DevOps. The mission statement was to be the show for people whose boss "read about DevOps in the in-flight magazine and now I'm supposed to do it." Matty learned from DevOps Cafe, then Food Fight and The Ship Show, and admits to still not always following what John Willis is talking about. Bridget points out that Matty is now name-dropping Willis and Damon Edwards after swearing off name-dropping in episode 1, and asks if Matty has met John Allspaw at Velocity. Matty hasn't been to Velocity yet.

                                                                      Trevor tells how they met at the new Azure meetup in Chicago, running into a "really loud, knowledgeable guy" who turned out to be Matty. Matty says follow-through has never been a strength, so a partner was needed for accountability, and that editing the first episode made Matty think "are you ever gonna let Trevor talk?"

                                                                      By the Numbers

                                                                      Matty cautions that podcast stats are guesswork, since they only count audio downloads, and that Matty halves them for personal use. The most downloaded and most viewed episode is 14, How to Eff Up DevOps, which Matty calls the Pete Cheslock effect (Bridget: "like the Colbert bump"). Older episodes keep growing, which Matty takes to mean new listeners are working through the back catalog. Matty says an episode gets about 2,500 to 3,000 audio downloads and 200 to 300 YouTube views, and that when the Minneapolis episode wasn't posted to YouTube, some people wrote asking what happened, since YouTube is how they follow the show. Every episode except the two live ones was recorded on a Google Hangout, "which is ironic, I guess, that our truly live episodes were not live streamed."

                                                                      Topic First or Guest First

                                                                      Matty admits part of the reason the podcast exists was to get Jez Humble on. After the tenth episode Matty tweeted, and Jez replied that the hard part isn't getting Jez on the show, it's getting Jez to shut up once on. Jez was episode 15. Trevor says the approach is topic first, then who can talk about it, and jokes that Matty likes famous names. Matty admits to being prone to the echo chamber, and says Trevor finds people Matty has never heard of, which is good since some listeners aren't in the DevOps community at all.

                                                                      Listener feedback shaped some episodes. People said the show was heavy on culture and light on tools, which led to the Git episode. Matty says the audience is causing an identity crisis, since it's meant to be the 101 show but people Matty knows listen. What matters more to Matty are the messages from people who say "I didn't even know where to start," or one who wrote that they'd drunk the Kool-Aid, their organization hadn't, and the show made them feel it could be done. That's "way more awesome" than knowing an expert listens.

                                                                      Favorite Moments

                                                                      The episode plays an audio supercut of the year's lines, including "tools are easy, people are what's tough." Bridget loves that in episode 1 Matty wouldn't read the John Vincent quote because Matty didn't want to say the word shit on air, and that by the Etsy episode "everybody is swearing like Chef employees." Matty says the swearing used to get a spring sound effect, until that stopped. Trevor says they originally avoided the explicit tag so as not to turn listeners away, and that Matty eventually decided that was a fruitless effort. Bridget also likes Trevor's line in episode 2 about never stopping learning, and Bridget's own MST3K-style response: "so you're like a shark?"

                                                                      Bridget's favorites are the Minneapolis episode with Patrick Debois, the conferences episode with Jason Dixon and Pete Cheslock, and hosting the enterprise one. Matty says that one was a comedy of errors, with hotel Wi-Fi, a guest in Budapest, and an Azure network incident that day. Trevor's favorites are episode 17 on asking for help and the Etsy story about telling a senior VP that the VP was being a jerk. Matty's is episode 14, partly for a running backchannel in a Google chat, and partly because Matty thinks a guest is different from a host, and it was fun to watch Nathen Harvey simply contribute. Matty regrets the awful audio from Matty's cheap headset.

                                                                      Behind the Mic

                                                                      For anyone thinking of starting a podcast, Matty says the cost ranges from nothing to thousands a month. They record on Google Hangouts on Air, which pushes to YouTube, then Matty pulls the audio and MP4, converts to AIFF, and either edits it personally or sends it to Mandy Moore (@therubyrep), who edits it and returns it via Dropbox, usually within 48 hours. The MP3s live on Amazon S3 and the site, which serves the RSS feed, runs on Azure. Sponsors pay for the post-production and microphones, and for the first few months it was funded out of Matty's pocket. Matty notes that when editing, everyone seems to say "um" more toward the end because boredom sets in, which is the argument for outside editing.

                                                                      The only thing stopping them from publishing more often than twice a month is time. A third co-host helps, since until Bridget neither could skip an episode, and Julian Dunn had to stand in at Minneapolis.

                                                                      The Retro Goes Missing

                                                                      Trevor says the original structure was a mock agile framework, which is why the closing segment is called checkouts and there was a retrospective at the start. They dropped the retro because they didn't have anything to say, and Matty says the current work can't be discussed, like customer visits. Matty now misses the retro, since shows like The Ship Show and Food Fight let you know what the hosts are up to, and asks listeners whether they care.

                                                                      Where the Show Fits

                                                                      Trevor says other podcasts rarely get a listen, and Matty says the show fits in the same place as before, alongside The Ship Show, whose hosts seem to hack Matty's Google Docs and do the show ideas first. Matty says the show is not what was envisioned, "and that's cool." Trevor says people have told Trevor they're glad to have an intro point, and a sysadmin friend prefers The Ship Show for the technical detail, and that for Trevor it's "nice to provide what I needed then."

                                                                      Bridget asks Matty whether the show has changed Matty's cynicism. Matty says the optimism was forced, starting when Matty told the team "let's just pretend it will work," and selling new ways of thinking has kept it up. The new Facebook search surfaced Matty's 2011 posts, where Matty was "a cynical asshole" about DevOps, including a coworker's comment that it was "great for tiny companies with no customers and no future." Bridget: "Yes, tiny companies with no customers and no future, like Etsy and Facebook."

                                                                      Both hosts changed jobs during the year, Matty to Chef and Trevor to 10th Magnitude, with about a week of overlap. Matty says the show improved networking, and at least once a month a prospective customer says "you do Arrested DevOps, right?" Trevor says 10th Magnitude was picked partly to learn from Matty, and that being on the show is now seen as an asset there.

                                                                      What's Coming in 2015

                                                                      Ducy is the field correspondent, and might pop into a Hangout from wherever Ducy happens to be. Planned topics, with the promise that they won't promise when, include a promise theory episode with someone who lives in Minneapolis, blameless postmortems, DevOps jobs, a Microsoft and PowerShell episode, Docker, eventually consistent distributed systems, and building teams. Matty says the previous longest episode was episode 2, at one hour and seven minutes, and that this one will beat it.

                                                                      Check-Outs

                                                                      Bridget saw the giant interactive Google Android billboard in Times Square that features DramaFever, and recommends Lara Hogan's Designing for Performance. Trevor recommends the Apollo guidance computer rewritten in JavaScript, the Human app now on Android, and the Netflix series Black Mirror. Matty recommends the Guardians of the Galaxy Honest Trailer, the expanded Facebook graph search, the @PeteChessBot Twitter bot, and the Accompli mail app.

                                                                      1 hr 14 min
                                                                    • The Database: The Elephant in the Room
                                                                      Why the Database Is Different

                                                                      Grant Fritchey, the self-described Scary DBA and an enterprise DBA for about 20 years, and Jonathan Hickford, a product manager, both from Redgate, explain why databases are the elephant in the room of continuous delivery. Grant's answer is data persistence. Releasing software is straightforward: "You take an old piece of software, you throw it away, put the new piece of software in, you're done." You could do the same with a database, "but then the phone starts ringing because the business is freaking out because there's no more data left in the system."

                                                                      Jonathan adds that data comes from several places, some generated by users and some configuration data from development, and often flows in the opposite direction from the code in a pipeline. A database may also be read by several applications and BI jobs, so a change has to be safe for all of them. Grant says data migrations for something like a new non-null column run against a live system and can cause downtime, and that part of the push for NoSQL comes from the difficulty of database deployments.

                                                                      Reviewing the Script the Day Before Is Too Late

                                                                      Grant hates the common DBA habit of looking at the script the day before it goes to production, which means the first run happens in production. Matty has seen DBAs who ran the developers' scripts without reading them and says that adds no value. Grant, self-deprecating, says the hard way taught that on Thursday you can't stop a Friday deployment: "What you have to stop is bad development," which means going to the scrums and standups and knowing what's happening six months, six weeks or six days before, not six hours.

                                                                      Jonathan says the core of continuous delivery is fast feedback, so a data problem should be found in development, and production deployment "should be boring." Grant says that line is getting stolen: people always ask how your plane flight was, and you don't want an exciting one, "because you don't want an exciting production deployment."

                                                                      Tooling and Source Control

                                                                      Grant says step one, though it sounds bad from a tool vendor, is tooling, since you can't get a database under source control by hand. You need differential scripts, alter scripts and migration scripts that move data, not create scripts rerun over and over. Grant's objection to ORM tools is that they only produce 1.0-style code, such as create table, when existing data has to persist. Jonathan says once the database is in source control you can repeat things, and test a refactoring against realistic data to find out it would take three hours over millions of rows. Grant adds that even knowing it takes three hours lets you schedule for it.

                                                                      The Myth of Rollback

                                                                      Matty says knowing you can't go back and are always going forward is a reason to test deployments continuously, since the old habit of testing your backout never happened anyway. Grant says only two rollbacks really work: restoring or undoing a snapshot right after the release, or redeploying forward. A script to roll back pieces is "such a lie because the data comes in over time and ain't nobody wants to get rid of it." So the deployment process itself has to fix mistakes.

                                                                      Grant says test environments should be as close to production as possible with cleaned data, for example without emails or personal details, and that sometimes it's fine to refresh QA when it burns down. Matty describes a former 3.5 terabyte data warehouse multiplied across more than 20 environments and being told there was no time to make test data. Jonathan says you don't need every environment to be a copy: one for data size, one for data complexity, and lighter ones for fast changes. Jonathan suggests testing sharding and replication earlier, and prefers "the short, fat pipeline rather than the long, thin pipeline."

                                                                      Blue-Green for a Database?

                                                                      A question from IRC asks how to blue-green a database change. Jonathan says it's possible with patterns, such as splitting a column by letting the application or data access layer handle both versions and migrating slowly at quiet times. Trevor asks about schema changes and Grant's answer is one word: badly. Grant has done it with triggers keeping two copies in sync, and calls it a "massive undertaking and fraught with horrific danger," like Indiana Jones. Jonathan says if a DBA says the database isn't the place to tackle this, they're probably right, and offers versioned schemas as an alternative.

                                                                      Matty adds that in agile shops the throwaway shim, like a bridge you won't need once the stream is gone, is a hard sell to a product owner with a user story waiting.

                                                                      How DBAs Feel About DevOps

                                                                      Grant says it's a mixed bag. The camp that upsets Grant most says everything's fine because they've always done it that way, and Matty's reply is "We have always been at war with East Asia." To reach them you have to document their pain: how long a deployment took, whether there was downtime or data loss, how long recovery took, and the same for the developers who hop through hoops before production, and take that to management. Others are ready, and Jonathan says at DevOpsDays and other conferences someone always asks in the Q&A "what are you doing about the data?"

                                                                      Grant describes the approach as that of a "completely lazy bastard" who takes advantage of what development teams have already worked out about source control, labeling and branching: "it's down to more labor than thought." Grant supported 10 development teams by automating builds and deployments, letting developers write their own T-SQL and reviewing for 15 to 20 minutes a day per team. Jonathan says the best visits Jonathan has made had a DBA, a developer and a sysadmin in the room.

                                                                      Common Tooling, Testing, and Meeting in the Middle

                                                                      Matty worries about silos created by tools, like data developers stuck on a different source control system. Jonathan says vendors are becoming more agnostic and that using one system lets you make atomic commits containing both database and application changes, which Jonathan calls probably a prerequisite for continuous delivery. Grant says without everything versioned together, people had to ask which database version went with which application.

                                                                      On testing, Jonathan mentions tSQLt and DBUnit, and Grant says you don't write a test for a table. What matters are tests on destructive changes, that a migration moved the same number of rows and the updated data looks as expected, for confidence and boring deployments. Jonathan adds monitoring after release. Grant says DBAs need to learn tools like TeamCity, Jenkins, Octopus and PowerShell, because "we can only automate all of that. So what is your job again?" and they need to meet developers in the middle, not stand "athwart the bridge stopping you from going to production." Matty says sysadmins need to get smarter about data too.

                                                                      1 hr 1 min
                                                                    • DevOps in the Enterprise
                                                                      DevOps Is DevOps Is DevOps

                                                                      This is Bridget Kromhout's first episode as official co-host. Bridget is an operations engineer at DramaFever with a past as "a BOFH at places many hundreds of times this size." The panel is Michael Ducy of Chef, Ross Clanton, who leads part of operations at Target, and Steve Pereira of MyPlanet in Toronto, who helps big companies get started with DevOps and Agile and runs the DevOps and OpenStack meetups there.

                                                                      Michael's position is "DevOps is DevOps is DevOps": the principles of continuous improvement are the same everywhere, and an enterprise just has more, sometimes political, nuances to dance around. Ross agrees, whether you're an enterprise, a startup or a unicorn, and adds that a large company like Target carries the technology debt it built up over the years, and a mindset shift across a large organization is a challenge that's "rewarding, it's frustrating, it's all kinds of things." Steve calls DevOpsDays "this magical camp where we all sort of get away from our jobs."

                                                                      DevOpsDays Inside Target

                                                                      Ross and Heather Mickman, a colleague, started an internal DevOpsDays at Target almost a year earlier, and have held three, quarterly, with a big annual one coming in February. The tagline is Connect, Share, Learn, and the aim is to get people across silos together by choice. The first drew 160 or 170 people, with Michael and an external speaker from Nordstrom, and Michael's talk "The Goat in the Silo" is still discussed by engineers there. Michael says enterprises don't realize they have enormous tech communities inside them that they haven't used, and that when Ross and Heather spoke at DevOpsDays Minneapolis about Target's transformation, it changed people's assumptions about the company.

                                                                      Do Enterprises Just Need the Automation Part?

                                                                      Matty says enterprise DevOps is often read as CAMS without the C, on the assumption that they don't need the hippie-dippy culture, just good automation and stats. Matty adds you can't cherry-pick, and that the traditional DevOps community talks about empathy so much this year because it feels it already has automation figured out. Ross says culture is the most critical part, since it amounts to mindset, and the rest follows as an outcome. Ross heard at the Enterprise DevOps Summit a mix of top-down and bottom-up approaches, and the theme that everyone is now shifting to optimize for speed.

                                                                      Steve says many of those companies' transformations took three to six years, and that culture is like steering a giant tanker, and reports hearing that many started because they were in a mess and had no other option. Steve hopes that won't continue, now that there's data. Michael disagrees: the body of knowledge was already there, and real change comes only from "oh my God, I have to," so "there's going to be a lot more bloodshed." Ross's own moment came from reading The Phoenix Project, with a sense of having lived its characters. Ross says the way to make DevOps stick in a big company is to find each person's pain points and help them past those, which is a localized and time-consuming approach, and not to inundate them with concepts like promise theory.

                                                                      The Conference Bridge Problem

                                                                      Bridget tells of meeting three people from Thomson Reuters at DevOpsDays Ghent, who mostly wanted to hear about team chat. Their pain point was incidents, where everyone sits on a conference bridge and each new arrival needs a summary, so they hear it "like 14 times in a row." Bridget showed them Slack, and two of them downloaded it at the table. Ross says persistent chat has created a lot of aha moments in some groups at Target too.

                                                                      Sharing, Open Source, and Auditors

                                                                      Matty says a few years ago it seemed only Netflix, Facebook and Etsy were doing DevOps, but the larger companies just weren't sharing, and some places told employees not to tell anyone how they did it. Ross says most enterprises don't fully embrace open source and rarely let people present or blog externally. Target's public tech blog was only three or four months old, and it's starting to contribute to open source, though "it's going to be a slog."

                                                                      Steve says one of the best threads at the Enterprise Summit was audit and compliance, with auditors on panels explaining how working with them within DevOps principles benefits a company. That matters, Steve says, because the fear that devs will run wild on production has kept many from taking DevOps seriously.

                                                                      Culture Isn't a Checklist

                                                                      Steve says the culture side is what everyone finds most intimidating, and that sharing is how someone at the bottom gets the conversation with leadership started. Bridget recalls Steve's Ignite about DevOps as a MacGuffin, and Michael's slide of an overly process-driven dystopia, where holding too rigidly to something like CAMS becomes dogma too. Michael says tools and even processes are easy, and that "the organizational management issues are the real challenge." Michael adds that the summit consolidated a lot of enterprise information in one place, and proposes an enterprise DevOps podcast. Ross is game, suggests "Arrested Enterprise DevOps," and Bridget looks forward to listening "from a safe distance."

                                                                      One Piece of Advice

                                                                      Michael's advice is that it's "probably going to piss a lot of people off" and you need to find allies. Ross says be very patient, and describes a Venn diagram that colleague Jason Walker draws on whiteboards, with three circles, collaboration, empathy and experiential learning, and DevOps at the center. Steve's is that sharing is how you start. Matty's is that you probably already do some of this under other names, and that with a transformation you should do a pilot and stack the deck for results. Trevor's is to advocate your problems and say what hurts, since you'll find everyone shares the same hurts. Bridget says you can't do everything at once, so if all you manage is getting ops into scrums, that's a start, or adding someone who knows what's going on to the change advisory board.

                                                                      Checkouts
                                                                      Ducy
                                                                      • The Goat Farm podcast
                                                                      • Ross
                                                                        • All the great talks from the Enterprise DevOps Summit, DOES14
                                                                        • Target Tech Blog (post on flashbuilds)
                                                                        • Steve
                                                                          • ALL THE TWEETS!
                                                                          • Debian snapshots
                                                                          • Amazon Lambda/CodeDeploy/Containers (in case it’s not covered in the AWS recap)
                                                                          • My recap doc of DOES14, pretty rough but shareable
                                                                          • Bridget
                                                                            • devopsdays.org - for a devops near you, or to be inspired by talks at past ones so you can hold one inside your org
                                                                            • Lots of great talks last week at DevOpsDays Vancouver, and one you should definitely watch is Stephanie Van Dyk, a Google SRE who worked on the healthcare.gov rescue.
                                                                            • Trevor
                                                                              • Worm Robot
                                                                              • Big Hero 6
                                                                              • Barbiefail
                                                                              • Matt
                                                                                • Yak Shaving Expert t-shirt
                                                                                • Customizing Chef by Jon Cowie
                                                                                • Have you tried turning it on and off again supercut from the IT Crowd
                                                                                • 47 min
                                                                                • git 101 with Emma Jane Westby

                                                                                  Microsoft is open-sourcing .NET and creating the CLR for Mac and Linux

                                                                                  .Net core5 is the new framework for 5, you can ship your own version of the app.

                                                                                  There is now a free version of Visual Studio — Visual Studio Community for open source developers and students.

                                                                                  In fact, it is all happening out in the open on github.

                                                                                  Emma Jane is a long time listener of ADO! She has been teaching version control for many years with specific emphasis on the communication behind version control in teams. She has since switched to distributed version control such as git. Her aim in her teachings is to create resources that make git ”less painful” than it currently is.

                                                                                  Distributed vs Centralized Version Control

                                                                                  Emma Jane:

                                                                                  • In distributed VC the DB that contains the changes, exists on the local system and I can have multiple connections to multiple DBs with other versions.
                                                                                  • Centralized is all in a single DB, locally.
                                                                                  • How is Distributed VC relevant to DevOps?

                                                                                    Matt: Many people hold the theory that you cannot have “The DevOps” without distributed version control. It implies communication through teams, so what is the validity of that statement.

                                                                                    Emma Jane:

                                                                                    • Git is not the only VC option out there, but it is the most popular currently.
                                                                                    • You need to assess your team and your project, along with the related expertise and community support, and go with the one that fits your needs.
                                                                                    • As soon as you say “can’t” someone will prove you wrong.
                                                                                    • Testing

                                                                                      Matt:

                                                                                      • The whole basis of git-flow is based on the fact that you can’t trust your contributors. Especially with open source. It sends a message that says “I don’t have to test my shit, because you’ll do it for me.”
                                                                                      • Emma Jane:

                                                                                        • If we are talking about testing, you need to have full coverage testing of whatever your product. Many testing frameworks allows for 99.9% accuracy on the tests, but that .1% causes you not to trust your tests. This makes it really hard to get reliable CI into the dev process.
                                                                                        • For that matter, Devs shouldn’t trust themselves when it comes to pushing code, you should always rely on testing because everyone is going to make mistakes, and humans might not catch them.
                                                                                        • Git allows you to have control over the pushed code.
                                                                                        • Trevor:

                                                                                          • There should be no permissions. All developers should have the same permissions and the flow should go through QA.
                                                                                          • How do you learn git?

                                                                                            Emma Jane:

                                                                                            • All kinds of people are interested in learning git. But mainly:

                                                                                              1. Someone who is on subversion and wants to change to git
                                                                                              2. Someone who has been told to use git, but they don't know how to run command line tools.
                                                                                              3. CTO or management types that know they want to use git, but they’re not really sure where to go from that decision.
                                                                                              4. In order to identify how your team will most efficiently use git, draw out your team flow and identify where efficiency is being blocked. Is re-basing causing problems? Is a PR sitting out there for too long? Use those as discussion points with your coworkers.

                                                                                              5. You cannot introduce creativity when you are just told to memorize commands.

                                                                                              6. Emma Jane:

                                                                                                - Use Interactive Add! It allows you to split up your diffs into different commits. So you don't end up committing a huge chunk of features that should most-definitely not be committed together.

                                                                                                How do I get set up?

                                                                                                Look for the right git-flow based on the type of deployments you are going to be using to release the software. Are your deployments feature based? or time based? How important is a rollback?

                                                                                                Your code should always be deployable in a CD framework. You are only rolling forward, you have one master branch, and feature branches, how can you have correct and fast CD if you have multiple branches before the CD process starts.

                                                                                                Your git setup should be directly related to your infrastructure. The git releases and flows of a team of 1 is going to be massively different than the git flow of a large team for a Could Provider.

                                                                                                Things you should know (about git):

                                                                                                Rebasing:

                                                                                                • Rebasing allows you to recombine how your commit chunks are strung together. It takes all the commits of a branch and
                                                                                                • Great for when you are adding too many commits.
                                                                                                • Git Bisect

                                                                                                  • You can take out commits individually, and assess if the commits are in a working state.
                                                                                                  • However, if you do not have full commits, for example, commits when you are just thinking about something, it will be much harder to assess the state of the commits individually.
                                                                                                  • Source of Emma's talk about Git: http://github.com/emmajane/gitforteams

                                                                                                    Post version: http://24ways.org/2013/git-for-grownups/
                                                                                                    Recording: http://prague2013.drupal.org/session/git-makes-me-angry-inside

                                                                                                    Emma's rant about storing the history of your project: http://gitforteams.com/resources/evolution-social-coding.html

                                                                                                    GitHub conversations: http://guides.github.com/introduction/flow/

                                                                                                    Check-outs
                                                                                                    Emma
                                                                                                    • Kaleidoscope mergetool -  because of image diffs
                                                                                                    • Sketch training for techies -  because of impostor syndrome for drawing
                                                                                                    • GNOME wins -  go open source!
                                                                                                    • Matt
                                                                                                      • Teamocil  - generator for tmux sessions
                                                                                                      • The NoPhone  - kickstarter for a “technology-free alternative to constant hand-to-phone contact”
                                                                                                      • “It’s not a promotion, it’s a career change”  - blog post by Lindsay Holmwood
                                                                                                      • Trevor:
                                                                                                        • Sandisk SSD
                                                                                                        • Android L: Inbox
                                                                                                        • Rosetta Probe
                                                                                                        • 59 min
                                                                                                        • managing systems in the cloud
                                                                                                          What "Cloud" Means Here

                                                                                                          The episode opens with the show reading its first glowing iTunes review, titled "Two Chimps and a Mic," which gives five stars because the reviewer "had no clue what they're talking about." Trevor's only correction is that there are two microphones. Then Tom Limoncelli, co-author of The Practice of Cloud System Administration, explains why the book exists. Tom's first book came out in 2001, and system administration has since shifted from help desk and machine administration toward service administration, including what Tom learned at Google.

                                                                                                          Tom says the word cloud "has been taken by marketing people and kind of destroyed" (the publisher wanted it in the title, so the book's first use of the word is to explain what they mean by it). What they mean is distributed computing, systems so large that work is divided among many machines, as opposed to the 1990s pattern of buying a bigger and bigger server, which stops scaling because machines only get so large. Trevor asks for horizontal versus vertical scaling, and Tom admits to keeping a cheat sheet on Tom's palm: vertical is one machine getting bigger, and horizontal is dividing the work across 8, 80 or 800.

                                                                                                          Data Over Hunches, and Analog Systems

                                                                                                          The first skill change moving to Google, Tom says, was that monitoring became far more important, because in a big distributed system no one person understands everything, so you have to instrument it and make decisions from data. Tom tells of the second or third week there, when Tom offered a solution from senior sysadmin experience and a very junior sysadmin pulled out a graph showing that the change Tom had made moved the line the other way, so they should do the opposite of what Tom said. Tom says it was humbling and "such a better way of doing system administration." The second change is that at large scale downtime is more visible, and you get good at sensing a sick system and fixing it before it becomes an outage, not at responding to outages.

                                                                                                          Matty asks about servers as interchangeable cogs, and Tom says "the components don't matter, the system matters." RAID is Tom's small-scale example of decoupling component failure from service failure. In a very large system, Tom says, computers aren't up or down but degraded: "Gmail is never up or down. It's running at 90% up or 80% up. It's an analog thing."

                                                                                                          Your Architecture Is Wrong, With Details

                                                                                                          Trevor's clients are finding their software is all-or-nothing, and Tom says that's why sysadmins need to be at the architectural level, not saying "we don't use that." Even in the most distributed system there's a dirty secret: "there's always that one thing that can't be replicated," like a lock server. When Matty cites Jez Humble's "your architecture is wrong," Tom says a sysadmin who says that gets asked how it should be done, and often doesn't know. So the first third of the book is distributed architecture explained for sysadmins, "like a boot camp for the architectural decisions," and the second part is operating large systems, with on-call and monitoring.

                                                                                                          The book's "surprise ending" is an assessment system, twelve questionnaires you rate from 1 to 5, and if you redo it monthly in a spreadsheet you get a heat map of improvement, going from red to green. It can be done per service, and rolled up across teams so a CIO can see where to move resources. At Stack Exchange, Tom says, the Director of Engineering used it to focus on the top-priority service.

                                                                                                          Twenty-Seven Handoffs

                                                                                                          To find where to start, Tom uses a lean analysis to find the bottleneck, or gets silos into a room to talk about the pain they cause each other. On one project the team analyzed one iteration and found 27 handoffs across 15 teams, and "no one had ever drawn a picture" of it. They sat down with each team and walked through the graph, and found many unnecessary dependencies and undocumented ones that explained delays. During a break, someone from one team cornered Tom, embarrassed to admit they'd never realized their delays caused five others. Tom says "an executive standing in front of everyone saying we have to break down the silos" has never broken one, but sitting down with another team builds empathy. In a sorting example, one team always handed the database down unsorted and every team after them spent two hours sorting it, when it could have been sorted once.

                                                                                                          At Stack Exchange, Tom adds, the first disaster recovery failover drill took 10 hours, then 5, and now 1 or 2. In one drill they filed about 30 bugs, and half were closed before it ended because developers watching operations do the steps built a button that did it in one click. Matty says you can't have "the empathy project"; Tom agrees the goal has to be improved service or uptime, and that empathy is the route. Tom's illustration is continuous integration, where the point isn't a two-hour build but the confidence to push constantly, and Tom compares software to a car company that warehoused every car for a year before selling it.

                                                                                                          Hipsters, Unicorns, and Damn Humans

                                                                                                          Matty says jargon can make the field intimidating, and mentions the blog post Matty wrote on hipster DevOps. Tom says technology always has the top 5% inventing things and someone has to push it to the other 95%, and is glad the DevOps community has stopped acting as if enterprises would never understand. On J. Paul Reed's closing talk at DevOpsDays Chicago about ending the unicorn talk, Matty says the unicorns themselves say "we're not unicorns," and that the concept of "the other" propagates the myth that DevOps only works at unicorn companies. Tom points to Todd Underwood's LISA talk debunking the idea that Google's problems are unique, and says every company has the biggest problem, "those damn humans." Tom cites The Forbin Project, where a computer concludes the real problem is people.

                                                                                                          Coping Mechanisms and "Bring Us the DevOps"

                                                                                                          Tom's time management book for sysadmins, which Matty has on Matty's shelf, is a list of coping mechanisms, since "no human actually has good time management skills." Tom kept it around 120 pages, with the top three tips in the first 20, because "always front-load."

                                                                                                          Asked what to do when the boss says bring us DevOps, Tom says to find out what they mean, since after a month you may hear "I don't see any difference" because they wanted something else. It's like "bring me the cloud," which to executives usually means getting a server in 15 minutes and not waiting 6 months, and to consumers means their photos are backed up. If they don't know, cynical Tom says fix your biggest pain point and call it DevOps, while non-cynical Tom says turn business priorities into measurable results, automate measuring them, and work on projects that move them.

                                                                                                          Check Outs
                                                                                                          Tom

                                                                                                          Stack Exchange is open sourcing its monitoring system called “Bosun”. Look for the presentation at Usenix LISA http://www.usenix.org/conference/lisa14/conference-program/presentation/brandt

                                                                                                          Matt

                                                                                                          pester busser for test kitchen by Jay Mundrawala (discussed in Matt Wrock’s post which I will link to because it is a long url)

                                                                                                          • Riffsy ios8 keyboard for animated gifs
                                                                                                          • Trevor
                                                                                                            • http://www.jeremymorgan.com/blog/programming/the-great-unicorn-hunt/
                                                                                                            • Borderlands the Pre-Sequel is out :D
                                                                                                            • 50 min

                                                                                                            About Arrested DevOps

                                                                                                            From the publisher's feed

                                                                                                            Arrested DevOps is the podcast that helps you achieve understanding, develop good practices, and operate your team and organization for maximum DevOps awesomeness.