Arrested DevOps

Arrested DevOps

By Matt Stratton, Trevor Hess, Jessica Kerr, and Bridget KromhoutTechnologyTech News
Download on the App Store

Arrested DevOps episodes

  • Making The DevOps Transition
    Two Routes to DevOps

    Jeanne Steinback, Director of Software Delivery at Rewards Network and formerly six years at Redbox, and Chris Andreen, Director of Software Development at the MacArthur Foundation, describe what implementing DevOps meant in each of their organizations. For Chris it's about technical debt and a history of doing everything by hand: "manual deployments, manual compilation, manual everything." He calls it pulling a thread on a sweater to see how far it goes, so the foundation can spend less effort on delivery and more on the development that supports its grantmaking.

    Jeanne's teams were already mature Agile teams with the low-hanging fruit picked. What surprised her was that a full-time DevOps person on the delivery team, who automated deployments and took them off the developers, "almost doubled" their velocity, since they were deploying two or three times a week and each deployment was expensive. Both cite consistency, since Chris says "nothing breaks the same way" and Jeanne says keeping environments in line used to be an enormous headache. Matty's summary is that it amounts to a release engineering practice under someone responsible for consistent releases.

    How You Know You're Happy

    Matty likes to start from success criteria: how do you know when you'll be happy? Jeanne's answer is "when your velocity goes up," since the only reason to have a development team is business value, and velocity is a hard statistic. Her team tracks cycle time from analysis to production, velocity including interruptions and bug fixes, bug counts, and deployments per two-week timebox, and baselines whichever area is a bottleneck before trying to improve it.

    Chris's metric is shifting the time spent cleaning up accidents toward work that adds value. His organization is unusual in that it doesn't take in money. "We actually don't take money in, so we do the opposite," and the goal is to be more efficient at giving it out and at tracking how productive the grants were.

    Resistance, and Fighting It With Data

    Chris hasn't met resistance, since he oversees both software and infrastructure and "the only resistance would come from me." Jeanne met plenty at Redbox, where there was no DevOps team and the need showed up by looking at bottlenecks. She won the argument with statistics, presenting where the time went and predicting a 20% improvement in deployment time, which turned out to be much higher. She used her own team as the guinea pig, and afterward DevOps was assigned to every delivery team, which are now "big proponents." At Rewards Network there was almost no resistance, because she runs all of delivery.

    Her advice is to accumulate data and make a logical presentation, not to say "I hear DevOps is really hot right now." Matty agrees, and says he's from Missouri: don't point him to a Gartner report, show him.

    Letting Go, Without Siloing

    What did developers have to change? Jeanne says "letting go of the responsibility for DevOps," because "it's hard to get developers to trust other people." Chris says the mindset he's still working through is the familiar one: things have been done this way for a long time and worked fine. Jeanne adds that once you're doing continuous delivery, "you almost can't live without DevOps."

    Trevor, a developer, asks whether that means siloing. Jeanne says no. The DevOps engineer sits next to the developers, and "stopped us from making so many stupid mistakes." Her point is that an Agile team calls each other out when someone's about to take a shortcut. Deployments no longer needed four, five or six people in a room, and became "hit the F5 button," because the DevOps engineers built the CI architecture together with the developers. Matty cites the John Vincent line again, and mentions the idea that it should now be called BizOps.

    Rebuilt Weekly, With No SSH

    Asked about his nirvana, Chris describes a plan to move all applications to the cloud, starting on Azure and later using AWS, and to rebuild infrastructure weekly from the latest server images so "we don't have to think about patching." There will be no SSH or RDP access, and servers will live for a week. He wants to keep things "as simple and as vanilla as possible," because if he can't just spin up a box and turn on the couple of features he needs, he's over-engineered the app. Databases will stay on-prem for now. He's heard the term server huggers and says they don't want to be that: "we're more than happy to just roll through servers like water."

    Matty points to an Ars Technica article on how Boeing merges its data centers with the Amazon and Microsoft clouds, and to the need to avoid provider-specific special snowflake features. He also credits Sascha Bates on the Mythbusters episode with the point that if the job is one nobody wants to do, like patching, ask why it isn't automated. He explains immutable infrastructure for listeners as the Netflix model, and quotes Steve Murawski that RDP is not an administrator tool. His joke about Chris keeping servers at arm's length is that "they're going to put them in therapy later in life because they didn't get enough love."

    Automate the Fifth Time

    For what worked less well, Jeanne says her team kept repeating tasks that could have been automated because everyone was rushed. They eventually put a big piece of paper on the wall, and when a task came up five times in two weeks, it became a candidate for automation. They did this for deployments, for development, and for interruptions where people kept asking the same question.

    Chris borrows the phrase "iteration zero," and says it's constant mistakes: they've been through TFS, Git and Subversion, and "restructured our repos 100 times," and he hasn't met the consultant who gets it right on every throw. On tools, Chris runs Subversion and Jenkins, and says he hasn't found anything Jenkins can't do. The one plugin besides Subversion is the Chuck Norris plugin, and a PowerShell library orchestrated by Jenkins spins up infrastructure and deploys code. Jeanne is moving from Subversion to Git and using Go and Mingle together, and lets her developers choose their own tools because she wants them happy.

    Make Work Fun

    That picks up a thread Matty has been on: his job isn't to make things suck less, it's to make work fun. He credits Jez Humble with the idea that DevOps and continuous delivery are about making it fun to be at work, which doesn't mean Nerf fights but enabling people to create things. In his words: "Patching servers is not fun. Troubleshooting a deployment script is not fun."

    Trevor prompts the question ops people always ask, which is what they'll do once it's automated. Matty says he heard it from his own sysadmins, and that John Allspaw said the same thing on the Etsy episode, with "are you kidding?" and a list of innovative work.

    Agenda
    • What problems were you trying to solve?
    • How did you (or will you) know when you are “happy”?
    • What resistance did you encounter?
    • What worked REALLY well?
    • What worked less well?
    • How Boeing merges its data centers with the Amazon and Microsoft clouds

      Check-Outs
      Chris
      • Bands in Town
      • Jeanne
        • Zapp! The Lightning of Empowerment: How to Improve Quality, Productivity, and Employee Satisfaction
        • The Night Circus
        • Trevor
          • OxDBE Jetbrains tools
          • Chrome dev tools - change css media state
          • Matt
            • Pragmatic Programming book tmux: Productive Mouse-Free Development
            • Burnout.io
            • 45 min
            • DevOps at Etsy: Not a Unicorn, Just a Sparkly Horse
              Not a Unicorn

              The episode opens with Trevor playing back Matty's promise in episode 1 not to name-drop John Allspaw, then noting that today's guests are the Etsy ops team, which Allspaw belongs to. Matty's defense is "It's possible." (Also in the news: Trevor has joined 10th Magnitude, and Matty is leaving it for Chef.) Jon Cowie, a senior ops engineer who has written much of Etsy's Chef tooling, "resolutely" maintains they aren't a unicorn, since that implies magical powers. They have servers that break, software with bugs and pages in the middle of the night. What's different, he says, is that "from the CEO down, IT and engineering is recognized as a core competency," and they've solved enough of the easy problems to be more proactive than reactive. His shortest version: "We trust our colleagues, and we communicate with them."

              With Cowie are Allspaw, who manages infrastructure and operations (roughly half of engineering), and web operations engineers Pete Bellisano and David Yurkiewicz.

              Nobody Says DevOps at Etsy

              Asked about the history, Allspaw says there wasn't a concerted effort, no lunch-and-learns about DevOps, and no marketing. It "still is very, very weird to hear the word DevOps" in the office, and if someone says it they're probably referencing something from the internet. What did change was that decisions had to recognize domain expertise, in a high-trust way, and beyond that the rule was "don't be an asshole."

              David calls DevOps a buzzword the industry has turned into a marketing term. Allspaw's take is that the word is "underspecified," like cloud, compliance or resilience. Nobody would say "let's add some more robustness" and treat it as an answer: "That would be the start of the conversation, not the end of the conversation." Matty agrees, calling the word a blessing and a curse, because it starts conversations and only becomes a problem if you stop there.

              Culture Isn't the Furniture

              Pete left IT about ten years ago, ran his own business for seven years, and came back to Etsy's corporate IT and then ops, and says the culture is "very nurturing." David says "you can't really buy culture," and that at earlier jobs the separation between sysadmins, developers and DBAs got in the way of ideas. Cowie, who works remotely from the UK, says the open-plan office isn't the reason: "I don't think the communication is a result of the layout of the furniture." The same cross-pollination happens with him over video chat and IRC, which David says is where most communication happens.

              Cowie also describes his own adjustment. He came from tiny startups where he was the only sysadmin and had "a very sort of B-O-F-H attitude," with the infrastructure his alone and someone screaming down the phone when it broke. At Etsy, he says, the idea was that with all the will in the world "you cannot defeat human error," so you ask what assumptions led to a mistake instead of yelling. Developers own their availability too, and there are developers on call who can be woken up.

              Making Misunderstandings Repairable

              Allspaw questions the premise that the culture is simply awesome. His view is that they don't indoctrinate people, they make misunderstandings "easily repairable," because bad cultures happen when practices drift and nobody can detect or fix it. He uses New Yorkers and drivers as an example: ask a New Yorker whether New Yorkers are terrible drivers and they'll say they're just drivers, while someone from Kansas City in the back of a cab will say something less polite.

              Cowie asks him to tell the NFS over WAN story. About a year earlier, Cowie had made an enthusiastic suggestion, and Allspaw, on a large conference call, called it the dumbest idea they could go with. His team told him afterward it hadn't come out the way he meant, and they were right. Cowie points out that Allspaw was then his boss's boss, and that in many companies you can't tell someone at that level they're being an asshole without being fired: "And now we're telling a story about it, and I still work here."

              Matty's view is that he probably wouldn't be fired, but the fear is enough. He once had a sysadmin ask what would happen when the CTO came running down the hall yelling to get the release out. Matty repeated it to the CTO, who said, "when have I ever done that?"

              Bare Metal, Small Deploys

              Cowie says Etsy's infrastructure is still "pretty much entirely bare metal," with S3 for some image storage and backups. The reason isn't aversion to the cloud but capacity planning, which Allspaw "literally wrote the book on." They run a monolithic PHP application on the LAMP stack, understand their seasonal traffic, and actually use the hardware they have, so moving to EC2 doesn't make commercial sense.

              They do a lot of continuous delivery. Chef manages everything under the application, while a separate tool called Deployinator ships the code, currently 50 or 60 times a day, mostly behind feature flags. The same ethos applies to Chef, where about 40 people have knife keys and make a couple hundred changes a month. Cowie says he expected the culture to fade as the company went from 250 or 300 people to nearly 500, and that instead, when his KnifeSpork workflow started slowing people down, the team simply sat down and changed it. He says he could take it on board without being defensive, and that there's "no defensiveness or empire building."

              Developers Who Deploy, and Ops Who Are Busy Anyway

              Allspaw introduced postmortems to Etsy, with the requirement that the room have diverse perspectives, including finance, legal, customer support, fraud detection and design. He says the effect of developers deploying their own code is that they care much more about how it runs in production: "we're not so secretly turning software developers who come work at Etsy into ops people." When people ask what operations does if developers deploy, his answer is "Are you joking?" and that there's an immense amount of proactive work. Matty says his own sysadmins asked the same thing about automation, and his answer was that they'd do cool stuff rather than copy files around.

              David describes designated ops people assigned to product teams whose job is to show them how to set up Nagios checks and Graphite graphs, and not to hold anything back. Cowie adds that a developer on the DevTools team asked to be in the on-call rotation and has been in it for about a year. Matty: "Somebody asked to be on call."

              Context in the Page, and What They Learned

              For the small problems they've solved, Cowie describes a teammate adding context to Nagios alerts: the page now includes a graph of the partitions, a Ganglia graph of disk space over time, and how far over the threshold it is, so at 4 in the morning you can tell whether you have to get out of bed. David's is host building: bare metal servers built in about 5 to 10 minutes each, so a new data center wouldn't mean building 100 servers by hand.

              Asked what they've learned, David says it's okay not to know everything, Pete says to keep asking questions, and Allspaw says every time he thinks a size barrier exists, he's proven wrong: "things fail at a certain size because you expect them to." Cowie's is "the consequence for failure is learning more stuff."

              • Episode 11: Etsy Examined - How the Best Do Their Business by Food Fight
              • BOFH (the Bastard Operator From Hell)
              • Check Outs
                Jon Cowie
                • dmg cookbook
                • rbenv cookbook
                • David Yurkiewicz
                  • Pushover.net
                  • Pete Bellisano

                    Buffalo Trace whiskey

                    John Allspaw

                    Lloyd Taylor "Hacking Your Organization"

                    Trevor
                    • Shortcut-Fu
                    • Agents of Shield
                    • Matt
                      • Vimium - Google Chrome extension which provides keyboard shortcuts for navigation and control in the spirit of the Vim editor
                      • Release! The Game
                      • 58 min
                      • Scaling the Application Mountains
                        ChefConf in a Paragraph

                        Matty had hoped to do a full ChefConf episode and didn't, so he points to The Ship Show's recap and calls bourbon and bacon the two best things about the conference. The remarkable part for him was Mark Russinovich of Microsoft keynoting to a room of 400-plus open source people, some of whom muttered "Here comes the sales pitch." Russinovich talked about Azure, said Titanfall runs every player on a VM in Azure, and showed Linux VMs being provisioned with Chef directly through Azure. The room was "pretty surprised and impressed," which made Matty feel good about the community.

                        Scaling Is Not Performance

                        Guests Steve Corona and Igor Papirov start with definitions. Steve co-founded Twitpic, wrote a book on scaling PHP, and now runs the API at Life360. Igor runs Paraleap Technologies, whose AzureWatch product monitors and scales Azure applications, and was previously chief architect at Restaurant.com. Igor separates scalability from performance: performance is fast or slow, while scalability is "maintaining the same performance over different peaks of usage." Steve agrees, and adds that scaling is "more infrastructure than it is code." What you prepare for, he says, is the stuff you never see at small scale, mostly timeouts and blocking.

                        Matty splits the load into predictable and not. Apartments.com has a seasonal business, and Cars.com knew a Super Bowl ad was coming. The Reddit hug of death, which Matty would have called the Slashdot effect, isn't predictable. Steve's Twitpic was unplanned success: a team of about seven, at most around 90 bare metal servers at SoftLayer, and petabytes of images on Amazon S3. He learned "by cutting my knuckles," rolling over in the middle of the night to restart Apache, and by crashing the site over and over. Many of the best practices, he says, "aren't published," and the top people just know them. He also wants predictability paired with configuration management, so knowing tomorrow is a big day doesn't mean building servers by hand.

                        Ninety Servers, All Doing Work

                        Igor says a core principle of scaling is getting every one of those 90 servers to do useful work, and that the first question is whether your design could use 200 or 500. He came up on N-tier Microsoft development, where the backend database won't scale to millions of people per hour. Steve says even at Life360, with about 100 AWS instances and plenty of scaling experts, some servers do less than they should, and distributing work evenly is "a very difficult thing to do right."

                        On vertical versus horizontal, Steve says everyone preaches scaling horizontally, but having more servers doesn't mean you scale horizontally; you have to plan for it. Igor adds that people scale relational databases vertically because they can't do otherwise. Matty says sysadmins were taught to start worrying at 70% utilization, because new hardware took eight weeks. Igor agrees that's enterprise legacy, and now "if you have 20 servers and they're all doing a little CPU, then you're throwing money away."

                        Build It Naively, Because You'll Rewrite It

                        Matty asks where the balance is between analysis paralysis and painting yourself into a corner, quoting the joke that "Rails app is up in 6 hours, it's down in 6 months." Steve says he has "the secret magic balance numbers" but isn't sharing them, and Matty offers 42.5 servers as the point to start caring. Steve's real answer is that the easiest way to build something is not to worry about scaling. You can still scale vertically a long way, since on AWS "you're one restart away from 500 gigs of memory," and he steers away from research paralysis by building "as naively as possible and then figure it out."

                        Igor agrees: a startup should get functionality out and listen to feedback, since the system will change several times before it becomes popular. Steve says you will rewrite your app, or rip pieces out of it. If you want cheap headroom, his suggestion is several small apps in the Unix philosophy. "It's going to happen."

                        Monitoring Is Not Diagnostics

                        For finding out what's wrong in production, Steve reaches for strace: with all the statsd and PagerDuty in the world, "you will never get the visibility that strace will give you when production is down right now." He says to learn exactly how your app runs end to end, and even read the source of the open source programs you use, because monitoring alone won't tell you. Matty says every good sysadmin still has a copy of What's Up Gold somewhere, and takes it as an example of DevOps having no demarc: quoting John Vincent, "never saying that's not my job," but also knowing when to escalate.

                        Igor's point is that monitoring tells you when things are broken, not why, so diagnostics is a different problem. He recommends monitoring from more than one system, since monitors go down too and each sees different things: AzureWatch knows the Azure infrastructure and New Relic knows the application code. Asked whether to pick one, his answer is "yes, you should use both," because tools that cost pennies per hour are worth it when you're making millions per hour.

                        No Batch Window, No Maintenance Window

                        Igor says large-scale systems can't have a batch window, since someone is always using the site, and nobody approves long downtime on a Sunday at 3 AM. Matty calls it a world of rolling maintenance windows, and jokes that maybe we should take outages so customers can have lives.

                        Steve offers the IRS site, which closes at 5 PM for EIN registration, and Matty explains it's probably an old CGI form emailing someone. Igor has a better one: a restaurant reservation service that dialed the restaurant and waited for a person to press a button. It's "pretty scalable," he says, since the waiting is sharded across hundreds of thousands of restaurants. Steve calls it "a whole new form of blocking I/O that I've never heard of before."

                        Parting Advice: Twelve Factors and the Single Brain

                        Steve recommends the twelve-factor guidelines from a Heroku engineer for building apps that scale more easily, with things like how to handle logs and not letting the server daemonize itself. Igor says the hardest layer to scale out is storage. A relational database is "a single brain" that can only scale up, so he'd push logic to the application layer and use object or NoSQL storage: "you're able to throw 1,000 application servers at the problem, and you can't really throw more than one SQL Server at the problem."

                        52 min
                      • Fast and Furious: Configuration Drift
                        Executable Documentation

                        Sean O'Meara of Chef, Chris Webber of Demand Media, and Steven Murawski of Stack Exchange each define configuration management a little differently. Matty calls it creating "a trusted target," since it's hard to automate deployment safely if you don't know what the systems look like. Chris, who works in Puppet, calls it "executable documentation that keeps reinforcing itself," which is accurate "assuming no one SSHs to it and fixes it." Steven stresses that it's an ongoing process, not a setup script, and that it only manages what you explicitly tell it to, so anything you change outside of it is invisible. Chris has another framing: treat servers as software objects that happen to have "really crappy APIs," with Chef or Puppet as the better API on top. Sean gives it the widest umbrella, covering any technique for managing configuration and its complexity.

                        Trevor, who's only recently started using it, likes that there's no lost knowledge, because everything done to a server is in the scripts. Steven says small-shop admins tell him they only have a handful of servers, until hardware fails and they're asking what registry key or service they had to tweak four years ago. Trevor's example is a coworker who objected to scripting the WebSockets setup when he could just click through Server Manager. The script came to two lines, and the coworker's reaction was "oh, that was way easier than I expected it to be."

                        Easier to Build a New One Than Fix a Broken One

                        On the business problems it solves, Matty says a trusted target enables more automation in the delivery process and shortens the feedback loop. Chris started out just wanting the same SSH config on every box, but what won people over wasn't Puppet itself. It was the stored configuration, which gave him a database of facts about every machine, and, for DBAs, machines that looked identical, down to the roughly 30 users Oracle needs before you can even install it. Matty adds that it kills "it worked on my laptop," and that laptops, outside the VMs on them, "should be used for email."

                        Steven's case is the one box out of four that behaves differently. Instead of two hours troubleshooting, reimage it in 20 minutes and pull its logs offline. Matty's rule: "it should always be easier to build a new environment than fix a broken one," and his nightmare is finding out months later that one of 300 web servers is different from the other 299. That leads to the Mark Burgess line, as Matty and Steven both recall it, that once someone logs in interactively you no longer know the state of the machine. Steven adds that even a login can trigger startup scripts or Group Policy.

                        Matty's own version: he was new to Puppet, saw broken images in QA, SSHed in and fixed a symlink by hand, and the developer confirmed it was working. Fifteen minutes later it was broken again, because a Puppet run had undone his fix.

                        Monoculture, Images, and Why Immutable Is the Wrong Word

                        Sean says a common way operations shops manage complexity is monoculture, being a Windows 2003 shop or a RHEL 5.8 shop until the next refresh. With configuration in code, moving from RHEL 5 to RHEL 6 means diffing the two and adjusting the policy, instead of boiling the ocean, and "it's a huge bummer when you want to run software and you can't because you're on a 10-year-old operating system." Steven agrees that easy replicas and version control make testing a new OS much less scary.

                        Sean dislikes the term immutable servers, because systems still need small changes. His examples are adding a machine to a load balancer pool and updating one firewall rule across a cluster: are you going to destroy every machine and push another 4 gigabytes down a pipe "when you need to change literally one line of config"? Matty explains the Netflix-style canary release that inspired the idea, and when Matty notes that most people aren't deploying ten times a day, Sean's answer is "they should be." Containers, Sean says, are really execution management and don't conflict with configuration management.

                        Matty credits Lucius from Food Fight with the term "leaden image" as opposed to a golden one, and a coworker suggested "brown and serve." Sean's position is that "images aren't bad. It's the loss of the resolution about the details of the image that's bad," as when teams with new VMware clusters hand-tuned a VM, snapshotted it and treated the snapshot as the artifact. Chris tells his developers that editing config on a box is like changing production code and never checking it in, or attaching a debugger to a running process and leaving no record. Sean's version: "why should my NTP configuration be any different than your HTML file? Stop it."

                        Same Test-and-Repair, Different Attitudes

                        Sean starts with what the tools share. They all derive from the idea CFEngine introduced, convergent test-and-repair operations: don't write the file if it's already right, don't install a package that's already there. A lot of people call that idempotent, he says, and he thinks they're wrong. Promise bundles, Puppet modules and Chef recipes are all named groups of those operators, and the tools differ in their philosophies on ordering and in language, with plenty of people detesting Ruby "with the blaze of a thousand hot suns."

                        Chris's short answer to Puppet or Chef is "just pick one and go with it." Puppet limits what you can do, so there's less chance to shoot yourself in the foot, and its dependency graph means a failing web stack doesn't keep SSH from being configured, which in Chef's run list could lock you out of the box. The flip side is that string manipulation, "90% of configuration management sometimes," is painful in Puppet. He used to steer old-school ops people to Puppet and developers to Chef, but isn't sure that still holds.

                        DSC and the Windows Problem

                        Steven says PowerShell Desired State Configuration "probably isn't ready for prime time yet as a standalone config management thing." Its key piece is the Local Configuration Manager, an agent on every Windows box reachable over WinRM, with a standard way to define and send a configuration, in the hope that Chef, Puppet and the rest will use it. He's been playing with its minimal built-in pieces at Stack Exchange, but wouldn't push anyone to be solely a DSC shop unless they were entirely Windows. Matty sees DSC as the agent that Chef can hand off to: a Chef recipe that runs a PowerShell script can only report an exit code, not confirm everything is as it should be. Steven adds that from Server 2012 on, PowerShell coverage grew from about 200 commands to about 2,400.

                        Chris says the hurdle on Windows is that Linux has a "package config service trifecta," and IIS on Windows doesn't; getting over it took "lots of bourbon." Matty spends his days configuring Windows with Chef: IIS is manageable, but a SQL Server cookbook works well for the single-exe Express edition and is much weaker for Standard or Enterprise, and it won't fix a changed configuration. Steven says official support for non-OS products is thin because DSC only shipped about six months earlier, and product teams respond to customer pressure. He curates the community repository at PowerShell.org, and tells Matty: "file an issue on GitHub. Seriously."

                        Where to Start

                        Chris points people to the Chef and Puppet Labs tutorials and to writing a little code, since "it's just like code." Sean's advice is to start small, model one class of machine, and resist using community cookbooks and modules at first: you'd need to be an expert in both the technology and the tool to read the errors. Writing your own makes the learning curve less steep. Chris likes starting with something identical everywhere, like SSH, and his first host entry points back to the Puppet master so a broken resolv.conf can't strand a box. Sean's advice on DNS: "Don't try to start DNS."

                        Steven suggests starting from your server deployment checklist: his Linux checklist had four items, his Windows one about 30, which is why building resources for it became his way in. Sean and Matty add that the most effective step is translating existing runbooks, not pasting them in. Sean tells of a customer with a runbook an inch and a half thick, and translating it exposed gaps: "install Apache. What version of Apache?" Chris asks vendors to publish a module alongside their docs, "in that format instead of English, which sucks for describing the actual state of the world."

                        Tools Are Cultural Artifacts

                        On the future, Chris says servers become software objects in "one gigantic development codebase." Sean thinks configuration management is "gonna eat the world" as developer culture does, since "tools are cultural artifacts." Steven expects the Windows space to be ripe for an explosion, and points out that executable documentation doesn't go stale like the inch-and-a-half binder. Trevor closes with what he remembers Matty saying earlier: the end of caring about individual machines, because you just spin one up and it stands up exactly as expected.

                        1 hr 6 min
                      • Fast and Furious: Configuration Drift - ADO9
                        One of the key technologies to help automate your DevOps environment is Configuration Management. There's a lot of chatter around what exactly this means, and how you can use it. Special panel guests Sean OMeara, Chris Webber, and Steven Murawksi join Matt and Trevor to talk about how Config Management can make your systems and stack more stable, predictable, and more fun to manage.
                        1 hr 6 min
                      • managing your mental stack
                        The Fire Hose, Push and Pull

                        Sasha Rosenbaum, a consultant at 10th Magnitude who grew up in Ukraine and spent about 11 years in Israel before Chicago, joins Matty and Trevor to talk about how to decide what's worth learning and how to learn it. Matty's method for the fire hose of information is to let it in without filtering: "my ears are wide open, my eyes are wide open," everything rattles around unprocessed, and if something sticks he digs deeper. He says he misses 90% of anything on Twitter, and the interesting links go into Pocket for later. Sasha's fire hose is Flipboard, plus a little Facebook and LinkedIn and the occasional tip from a colleague. Trevor's starts with Reddit and the meetups he goes to, and then he takes what sticks to Google. Sasha's name for the pairing is "a pull request and a push request."

                        Matty found configuration management, now a big part of his job, through a small reference in the Continuous Delivery book. Trevor opens with a question the episode keeps circling: "How do you find the best balance between Renaissance man and specialist?"

                        Solve It, Then Write It Down

                        Trevor learns by doing and calls himself "very bad at conceptual learning." Matty agrees, and adds that books make it easy to overestimate what a tool can do, because you never had to make it "jump through a hoop." Sasha's version is that books describe the good stuff, and you only find the one thing a tool doesn't do once you're committed to it. Her plan, which she says she hasn't done yet, is to write a blog post after solving a problem, since "you learn, you solve, and then you teach." Trevor treats his posts as "a secondary memory file." Matty's SharePoint search post still gets traffic, which he says is "relevant to the internet, it's not relevant to me, thank God."

                        Sasha tells of a friend who handed her a PDF on an unfamiliar algorithm and asked her to say when she understood it. After five minutes she thought she did, and then she was asked to teach it and found she didn't. Trevor says the same thing happens when he explains one solution to five people: the first time he's still thinking it through, and by the last he thinks "I actually truly understand" the solution, which is also why he likes pair programming.

                        Podcasts Get Authoritative

                        Matty reads out a friend's line that listening to podcasts is "a terribly inefficient way of transferring actual information," like "asking Grandpa Simpson for directions." His own experience is split. An audio-only podcast on writing Rails code didn't work for him, but he absorbs concepts from Food Fight or DevOps Cafe better than from docs, and a line he heard there, that a role should be an alias of a run list, stuck because he can hear it in his head. He's caught himself quoting podcast guests to clients as though they were authoritative, and Trevor's answer is "why shouldn't it be?"

                        Sasha has a 40-minute bus ride each way, gets carsick trying to read, and found podcasts fill the time, though she wants to argue back and can't. She thinks of it as taking part in a discussion instead of passively consuming a lecture, and says it "feels like having a discussion with your friends about things you care about." Trevor has the failure mode: he was listening to the Ship Show and started to join the conversation, before realizing "this isn't live," and "there's not real people here right now except the actual people on the train."

                        Nobody Gives You a Training Day

                        Matty says the rare job gives you time to just go learn, and that conferences work for him as a way to make connections, not to learn: "I learn nothing at a conference directly." Trevor's office tries a weekly brown bag, but after a day of bashing his head against a problem he either keeps working or veges out, so he reads new material on the train in the morning before his brain is overloaded. Sasha says a weekend of planned learning never happens, and that "go learn some random thing" wouldn't work for her, while a proof of concept for a specific problem would have her working after hours.

                        Sasha's other tip is switching context: read about business or networking or neural pathways instead of more technology. Matty agrees and cites The Goal as reading on what your company is actually for. His own rule is about an hour a week on something outside his focus, like learning Ruby without any connection to Chef, and he then notes that DevOps hasn't come up until three-quarters of the way in. Trevor says they'd mentioned it "at least 4 times," and Matty says all of them were in the first three sentences.

                        A Little Bit of the Other Side

                        Sasha describes her first reaction when 10th Magnitude added infrastructure automation people: "I am a developer. I don't want to do infrastructure." It changed when Matty joined and they started talking about DevOps instead, which she found made "total sense" because she could get exposed to a little of the adjacent side, set up a CI or automate a delivery, and still have an expert to ask. Matty's own version is learning enough about development to know how to talk to developers.

                        Trevor says he leans toward being the Renaissance type because he finds both the developer problems and the ops problems fun. Matty points out that this sounds like every problem is fun to him, and Trevor agrees he can't name one he enjoys more. Sasha adds that a calculus problem would be another matter, and Trevor says "calculus can stay somewhere else."

                        Pocket, Tabs, and the Pomodoro Challenge

                        Matty's workflow starts with a quick triage while he's on a break: links from Twitter and Flowdock go to Pocket, which is "good, bad, good, bad," so he can do it standing on the L platform. When he has 20 or 30 minutes, he goes through Pocket and reads for real. Trevor leaves Reddit links open in tabs, and jokes that he reads them once a month on Tuesdays, conveniently right before the updates force the machine to restart. Matty's name for it is "Patch and Read Tuesdays."

                        Sasha brings up the Pomodoro technique, since her concentration only lasts about an hour and a half. She also summarizes an article on productivity whose first tips were to take more breaks and naps, and whose third was to spend more time in nature: "I'm going to go find a tree in Chicago and stand next to it." Trevor turns it into the episode's challenge, which is for listeners to try Pomodoro, and use the breaks to catch up on reading, and report back to @ArrestedDevOps.

                        • Matt's popular blog post - Configuring SharePoint 2010 Search in a one-way trust scenario
                        • Food Fight - Episode 36: Roles, Environments, Attributes, and Data Bags
                        • The Goal by Elliot Goldratt
                        • Flipboard
                        • Pocket
                        • Pomodoro
                        • Vitamin-R
                        • Check-Outs
                          Matt
                          • Vagrant 1.5
                          • QuizUp
                          • Trevor
                            • .NET Fiddle
                            • Air Disasters
                            • One-Wipe Charlies
                            • Sasha
                              • Visual Studio Web Essentials 2013
                              • Self Promotion for Introverts
                              • The Little Prince
                              • 1 hr 2 min
                              • Managing Your Mental Stack - ADO8
                                It's information overload these days - how can a technology professional manage to keep up with everything that is new and exciting in the world of DevOps? What are the best methods for absorbing content and developing skills? Just how useful are podcasts, anyway? Matt, Trevor, and special guest Sasha Rosenbaum discuss these topics and their own personal strategies and challenges with keeping their brains from segfaulting.
                                1 hr 2 min
                              • all together now
                                It Has to Be Substantially Better Than Email

                                Matty's retro is that his team has been piloting Flowdock for talking to on-site consultants, and that Hubot is now installed there, mostly to post memes and Breaking Bad quotes. That sets up the topic: collaboration, with guests Angela Dugan and Todd Vernon. Angela manages the ALM practice at Polaris Solutions in Chicago and spent about five years as an evangelist at Microsoft. Todd started out writing software for X-planes at NASA, went on to found Raindance Communications and Lijit Networks, and now runs VictorOps.

                                Matty says he hates email "with the fury of a thousand suns," and asks how you get people to switch. Todd's answer is that a new tool "has to be substantially better." He thinks collaboration works best when it's vertical instead of horizontal: HipChat is horizontal, so a salesperson, a DevOps person and a developer can all use it, but you lose the value of a platform that combines the data and the dialogue. Angela wants a natural fit. Email reminders "would kind of make you angry" because they yank you out of what you're doing, so she looks for a tool that plugs into what people already live in, with low friction and a phone version. Todd adds that for DevOps it has to be nearly mobile-first, because you might need to communicate on weekends, in the middle of the night, or "if you're skiing in the mountains."

                                One Stream, Infinitely Filterable

                                Todd says his company argued internally over whether the right metaphor was chat or a Twitter-like stream, and chose the stream, so that the alert feed and other events sit alongside what people are saying. Chat, he says, tends to make people talk about the problem and leaves no room for updates and data.

                                Angela says information isn't collaboration if it's an undifferentiated deluge: the DevOps person shouldn't have to filter everything to find what matters. Matty tells of an application throwing errors nobody could interpret until Splunk dashboards put the number, 14,000 errors a minute, on a big screen the whole company walked past. "Very quickly it got looked at." Todd's ideal is a single timeline, "the black box recorder of the whole business," infinitely filterable, because otherwise a postmortem means pulling the story from seven different sources.

                                Use Your Feet

                                Angela cautions that dashboards without context invite people to "do very evil things with data," such as judging a team's value by bug counts and severities, so teams still need to meet face to face, and product owners and stakeholders should be invited to retrospectives. Her closing tip for the episode is the same: "use your feet." Email and IM become a crutch, tone gets lost, and if you can't be in the room, get on a hangout.

                                A Common Language

                                Todd asks whether writing software has become more social than it was a decade ago, and whether codifying deployment has given developers and ops a common language. Matty agrees on both. He came up as "the grizzled old sysadmin that sat in my silo and bitched about those cowboy developers," and says infrastructure work has become more like development, which is where it started, since the earliest sysadmins had to write their own tools. He thinks the meld is easier for developers, since ops is the one adopting practices like version control and test-driven development, "hell, testing at all." Developers building services have to care about operations in ways they didn't before.

                                Angela describes consulting years when "we were the ones pushing things to production," then the shift at Microsoft between 2005 and 2011, when infrastructure people started showing up to her ALM meetings to ask about source control and deployment tools. In her experience, companies that treat software as a social activity are "far less dysfunctional," with less contention between teams and less sandbagging on estimates. Trevor gives the counterexample: in a conversation about preparing the development environments so they'd fit into the client's production environment, the response was a flat "oh, that's up to them."

                                Will There Still Be a DevOps Group?

                                Todd asks whether a separate DevOps group is a temporary situation, and whether in ten years everyone will be in it. Matty says DevOps is "not a tool, title, or team. It's a philosophy," and that a delivery team can and should be cross-functional. As he told sysadmin teams he managed, not everyone should be able to do everything, but anyone should be able to do most things. Trevor mentions he'd just been elected stack lead of DevOps for CI, Azure and AWS. Matty's response is that it's "more work for no more pay," and Todd agrees.

                                Matty calls the cross-functional team "more aspirational than executed." Todd, who sees how many teams use his product, pushes back: old-school companies deploy it to an ops group of 20 or 30 seats, while new-school ones deploy it to hundreds, and at the most progressive, "everyone has pager duty." Todd's verdict is "I think we're both right," and Matty adds that more places see the value without knowing how to get there, "and that's why people like me have a job."

                                Age, Fear, and the Old-School Ops Person

                                Todd asks whether resistance is generational. Angela has seen a mix: sometimes the people who've been there 30 years are the most fed up and ready for change, and she thinks the fear is about confidence and personality, and about whether management will accept that "it might be ugly for a few months." Trevor thinks some resist because they fear extra responsibility, expecting ten more hours a week "because now I'm a DevOps person." Matty says it's experience, not age, and that if you've been somewhere 20 years, "the ship probably isn't sinking."

                                Todd has empathy for old-school ops, who have to learn to code and get no credit for keeping the machine running. Matty agrees, and says his profile on his last employer's Yammer read "you have no idea what my team does, and that's a good thing." His joke is that ops "doesn't have a dog to kick." He argues that continuous delivery is really built around stability. Todd adds that it isn't less work for ops, but the problems are smaller, and "I've never known one of these teams to get smaller." Matty recalls telling a team he managed that he was paying them "an awful lot of money to copy files around" and would rather they innovate.

                                Stand-Ups for the Waterfall Crowd

                                Angela says the best practices of Agile are really just best practices in software development, and she has convinced some very waterfall customers to hold a daily stand-up anyway: "I don't care if you only deliver once every 4 months, you should still be meeting every day for at least 15 minutes." Her reasoning is that a company releasing to the public every six months can still work in small chunks internally. She wants testers, BAs and PMs doing the same, and across products too.

                                Asked for one tip each, Matty says: "No email. Stop using email. Find something else." Angela's is to use your feet. Todd's is to quantify your job in terms of value to the business, since most teams he talks to don't know the value of downtime, and so can't argue for the tools they need. Trevor's is to know how to talk to everyone on your team: if the CEO walks by and asks about your project, be able to describe it in a way that makes sense to them.

                                Check-Outs
                                Angela
                                • Drive: The Surprising Truth About What Motivates Us
                                • SockDreams - @SockDreams and www.sockdreams.com
                                • Todd
                                  • Best BBQ chicken receipt in the world on http://www.thepauperedchef.com/ : http://bit.ly/1chcmHQ
                                  • Aberdeen report on DataCenter Downtime: How Much Does it Really Cost, free to download here: http://bit.ly/Mpd2E0
                                  • Matt
                                    • Meez - Setup tool for Chef authoring
                                    • Downtown Chicago Azure Meetup - Feb 27, 2013
                                    • 1 hr 1 min
                                    • All Together Now - ADO7
                                      Angela Dugan of Polaris Solutions and Todd Vernon, CEO & co-founder of VictorOps, join the ADO crew to chat about the challenges of collaborating in a cross-functional team. How can tools help facilitate communication among developers, testers, and operations? What are some of the best practices to keep in mind? And, of course, there just might be some "horror stories" of communication gone horribly wrong.
                                      1 hr 1 min
                                    • DevOps mythbusters
                                      An Aspiration, Not a Finish Line

                                      Damon Edwards, who runs DTO Solutions and SimplifyOps and co-hosts DevOps Cafe, and Sascha Bates, who works for Chef and co-hosts The Ship Show, take a list of beliefs Matty and Trevor compiled from clients and rule each one a myth or not. The first is that you're either DevOps or you're not. Sascha's answer is "I don't think anything is absolute, ever." Damon defines DevOps as looking at everything from a business idea to a customer outcome in production and removing the bottlenecks in between, which means you never finish, so it's "more of an aspiration" than a state.

                                      On whether DevOps is only for startups and web companies, Damon says any business that runs on software can benefit, and web companies simply got there first because they have no legacy and "this is their factory floor." Sascha thinks it's "more likely more valuable in big companies," since that's where the silos are. When enterprises push back with "we can't do things exactly like Facebook does," Sascha's summary is "My special snowflake is too special." Matty says a client once pointed out he was the second person to compare them to Netflix.

                                      Behind a lot of that is the way we've always done it. Sascha says the reason is often one that "isn't even valid anymore, and often was implemented by somebody who isn't even there anymore." Damon calls them hallway requirements. Matty tells clients that "nobody does things because they're dumb," and Trevor adds the follow-up: "Is that reason still valid?" Matty: "Doesn't mean we can't change it."

                                      Silos Grow With the Company

                                      Does DevOps scale? Damon's answer is that it's the wrong question. In a five-person garage startup there are no silos "because everyone's on the same team," so there are no DevOps problems. The problems arrive as the company grows and the handoffs go bad, which makes DevOps "more and more essential as you scale." Unlike scaling Agile, he says, it isn't a methodology: "It's a mindset. It's a goal. It's a set of practices." Sascha offers floating experts and cross-functional generalists as ways to adapt it, and says there's "no magic sparkle dust" and no way to make people do it.

                                      Damon adds that people are hard in every field. Manufacturing lived through these problems, and so did "organizing armies," and the industry shouldn't think it's made of special snowflakes.

                                      Agile Without Scrum

                                      Can you do DevOps without being Agile? Sascha says you probably won't be able to help it, since DevOps, Agile and Kanban share a goal of fast feedback loops. Damon notes that a lot of people equate Agile with Scrum, and if that's the definition, then no, one doesn't depend on the other. If you mean fast feedback, small batches and the lean principles underneath, "you really can't escape getting there."

                                      His evidence is HP's printer firmware division, which worked out a continuous-delivery-style process by thinking through its own quality and throughput problems, and only afterward read the books and found a match he calls "pretty much a one-to-one." Matty says the first DevOps Cafe episode he ever listened to, with Jez Humble, had him yelling "that won't work" at the car radio until Humble brought up the HP book: "oh, okay. Never mind."

                                      Matty then has trouble stating his own answer. He's worked on a DevOps-style team that wasn't Agile, because the product didn't suit small feedback loops, and it still helped to erase the line between sysadmins and developers. He lands on a verdict of "No, it's not a myth. It is a myth," Trevor offers "It's plausible," and Sascha says the two are "correlation without causation in some ways, and they're apples and oranges in other ways." Matty's excuse is that "there's too many negatives in this thing."

                                      The DevOps Team and the Canary

                                      Sascha's view of a team called DevOps is that it's usually the DevTools team, and "creating a siloed team to break down silos is really highly ironic." Damon's version starts from the release function, which is usually where things go wrong first and is the canary in the coal mine. When the canary falls over, "everybody says, well, we need a stronger canary." Then the release team gets renamed the DevOps team, and the silo has just changed its label. Sascha adds that release engineers don't like being rebranded, and "you're just stamping a label on them." Matty has also seen new teams built to go around ops, which makes things worse, and Damon says the current name for that is cloud operations.

                                      Damon does think a team can work in two forms. One is a Toyota-style chief engineer, or value stream manager, who owns the flow of work end to end, which he says is more of an architecture job. The other is operations as a service, along the lines of a Netflix talk on metrics, where the monitoring group doesn't take tickets but builds services, APIs and libraries the rest of the company consumes. Put the DevOps title on it, though, and "that's the DevOps team's problem" is what you'll hear.

                                      Sascha had a client who said "we have a DevOps team. We're not really sure what they do." Matty notes that in Chicago a DevOps engineer job listing is a sysadmin job, Damon says the title used to be a helpful hiring signal, and Trevor says he sees it used for developer tools like Git and CI. Sascha: "The words DevOps tools make me cry."

                                      One Gold Bar in a Pile of Straw

                                      On "we can't do DevOps because we need separation of duties," Sascha reaches for a word considerably ruder than myth, and Trevor corrects her: "It's called a myth." Her argument is that better configuration management lets you isolate the parts that need separation, such as customer data, without locking everything else in Fort Knox "because you've got one gold bar in your giant pile of straw." Matty says a lot of these myths could be rebranded as excuses, compliance among them, and Trevor says it was hard not to use that word when writing the list.

                                      Damon says people often can't say why they're doing it beyond an external force like PCI, SOX or HIPAA. His point is that the control is weak if Matty commits the code and Trevor deploys it as change one of 500 with no real validation. A continuous delivery pipeline that everyone adds tests and checks to works like an immune system, he argues, and those environments can be more secure and compliant than a system with one role that commits and another that deploys. He also says that for most of those requirements, separation of concerns isn't something anyone is actually checking for.

                                      Operations as a Service

                                      Does DevOps mean developers do the ops work? Sascha says it depends on what the day-to-day work is. If it's tedious, it should be automated, developers should care for the health of their apps, and "the ops aren't babysitters." Damon's model is operations as a service. Operations has historically been ticket-driven and a bottleneck because developers vastly outnumber it, so turn deployment, environment management, restarts and diagnostics into self-service the rest of the organization can use, which frees operations to teach others and improve infrastructure. Matty's condensed version is that they don't do the work, they facilitate it.

                                      Damon cites John Willis's name for the goal, the 80/20 flop: operations now spends 80 to 90 percent of its time "in the muck," and that should be flipped.

                                      Root Access, Children, and Visibility

                                      On developers getting admin access in production, Sascha says that "if you treat developers like children, they're always going to be children," and asks why anyone is logging into production given today's tooling. Developers, she says, don't want to hand ops things that break, and given tools to monitor their apps they'll want to use them. Matty repeats a quote he attributes to Mark Burgess, that every time someone logs onto a system interactively, they compromise everybody's understanding of that system, and admits that when he planned to audit his team's interactive logins as part of their reviews, he was the worst offender. Sascha ties it to blameless culture: "if you make it a crime to make mistakes, people are not going to own up to mistakes."

                                      Damon pushes back on both sides. Nobody needs root, but it's a little childish when developers demand it with "trust me." The real question is whether they have the control and visibility to do their jobs, and his formula is to centralize standards and decentralize control. Giving everyone root in a large organization is, he says, "a little bit naive."

                                      Windows Shops, and Tools That Enable But Don't Promote

                                      Matty rewords the open source myth into a less absolute version: DevOps works better with open source tools, as opposed to not working at all in a Microsoft shop, and he notes that many of his consulting clients are Microsoft shops. Damon says the difference is small composable tools that are API-driven, against big integrated point-and-click stacks, so it's about ease of integration more than licensing. Matty agrees Linux is easier in practice: Test Kitchen's answer on Windows is that "it's not that we don't want to support Windows, we just don't know how to do it." Damon points to Jeffrey Snover's DevOps Cafe episode on Microsoft's GUI-centered history and the new PowerShell work, and Matty adds that System Center Configuration Manager is "a big database" and you can't version data.

                                      Do tools promote cultural change? Sascha says they can enable it, but "promote is too strong a word," since without the cultural pieces a tool only exacerbates friction. Damon says people like talking about tools and think they're more self-aware about culture than they are, so they recreate their old broken world in the new tools. And "nobody ever gets fired for a successful tool implementation": the statement of work is complete, the tool works, and the person who installed Chef gets a promotion whether or not anything improved. Matty: "You got to add Chef to your resume."

                                      A Management Problem, and the Fish Market

                                      For the last myth, that DevOps doesn't work, Sascha says it depends what you mean, and she's reached the point of not wanting to use the word at all: work on communication and collaboration and call it whatever gets it done. Matty says his company has to use the word because customers ask for it, but "we don't sell it as fairy dust."

                                      Damon says DevOps is a management problem, and management has to align the organization around a goal. Tell developers to carry a pager with no context and the response is "Screw you. Who are you and what are you telling me all this stuff for?" Sascha adds that the message has to be relevant, and tells of a company she worked for where the Seattle fish market video came down from the top as a lesson in cooperation, and everyone below shrugged it off. "Don't try to peanut butter over with kumbaya."

                                      Damon's Ford example is that the CFO could probably explain how a car is made, and asks how many technology companies have executives who could say how their software gets made. IT has played the high priest, he says, and business leaders have let it: "It's the Jedi droid trick."

                                      Myths!

                                      The intro

                                      • You’re either DevOps or you’re not
                                      • The Company

                                        Management Beliefs

                                        • DevOps only works for startups or web companies.
                                        • DevOps doesn’t scale.
                                        • You can’t do DevOps without being Agile
                                        • Business Semantics

                                          • Shops practicing DevOps should have a DevOps team.
                                          • We can’t do DevOps because we need separation of duties.
                                          • The Team

                                            Operations Assumptions

                                            • DevOps means “developers do operations work”
                                            • A DevOps is a sysadmin that uses config mgmt.
                                            • DevOps is about hiring sysadmins who code.
                                            • Developer Expectations

                                              • DevOps means developers get admin access in production
                                              • Developers cannot be trusted.
                                              • The Tools

                                                • DevOps only works with Open Source tools and operating systems (i.e., I can’t do DevOps in a Microsoft shop)
                                                • The tools promote the DevOps cultural change.
                                                • The wrap up

                                                  • DevOps doesn’t work.
                                                  • Reference Links
                                                    • A Practical Approach to Large-Scale Agile Development: How HP Transformed LaserJet FutureSmart Firmware
                                                    • DevOps Cafe Episode 33 - Jez Humble
                                                    • Release Engineering Tools at Netflix - The Ship Show
                                                    • Keep Calm and PROD On - The Ship Show
                                                    • DevOps Cafe Episode 36 - Jeffrey Snover
                                                    • There's No Such Thing As A DevOps Team - ContinuousDelivery.com
                                                    • Check-Outs
                                                      Matt
                                                      • gitdrunk.com
                                                      • Downtown Chicago Azure Meetup - Feb 27, 2013
                                                      • Trevor
                                                        • Marvel Comics API
                                                        • Sascha
                                                          • DevOps meetup in Minneapolis
                                                          • Everyone should submit a talk for a conference!
                                                          • Damon
                                                            • QCon Conference - DevOps track in London in March
                                                            • Rundeck 2.0 just released
                                                            • 1 hr 7 min

                                                            About Arrested DevOps

                                                            From the publisher's feed

                                                            Arrested DevOps is the podcast that helps you achieve understanding, develop good practices, and operate your team and organization for maximum DevOps awesomeness.