SaaS for Developers

SaaS for Developers

By GwenTechnology
Download on the App Store

SaaS for Developers episodes

  • Cloudflare: Performance isolation in multi-tenant DB

    Cloudflare is no longer "just" a CDN serving 55M HTTP requests/sec; they now offer a wide range of cloud services on the edge. These services run on a data layer with 15 Postgres clusters running hundreds of databases.

    Vignesh Ravichandran, engineering manager of Cloudflare's database team, joined us to discuss the challenges of running this large-scale multi-tenant environment - dealing with network partitions, noisy neighbors, floods of connections, and even global warming.
    We talk about the importance of having a good toolset, of practicing incidents, and of internally advocating database best practices to a large engineering organization.
    The blog: https://blog.cloudflare.com/performan...
    The Scale presentation: https://www.socallinuxexpo.org/scale/...

    37 min
  • Real-time Data Infrastructure - At Uber and Beyond

    Pinot team at Uber wrote an excellent paper about the real-time analytics platform they built. Chinmay, formerly a principal engineer at Uber and now head of product at StarTree, joined me for a conversation.

    We discussed the challenges they encountered at Uber, the solutions they came up with, the platform they built, and how to best apply their experience to companies much smaller than Uber.
    The paper: https://arxiv.org/pdf/2104.00087.pdf

    39 min
  • Scalable Multi-tenant Platforms at Loom and at Times

    Shayon wrote a great blog post on the guiding principles he and his team at Loom used to guide them as they evolved Loom's data platform through a period of hypergrowth. I invited Shayon to the show to discuss the challenges he encountered and how he solved them - and I learned that he is now at Tines - solving an entirely new set of challenges with a very different set of solutions.

    We discuss the fundamental principles that help Shayon make critical decisions when building and scaling his platforms. We also discuss what happens when JSON columns become too large, how to prioritize performance improvement projects, what to do when a tenant gets noisy, and how his team performs zero downtime upgrades.
    Shayon's blog: https://www.loom.com/blog/scale-engin...
    Shayon's personal site: https://www.shayon.dev/about/

    48 min
  • Building SaaS on Kafka Streams

    Colt McNealy is re-imagining the future of microservices orchestration and he decided to build it entirely on Kafka Streams.

    In this conversation we discuss how Kafka Streams provides the low latency, reliability, availability and elasticity that is needed for the next generation of microservices orchestration. Colt also shares the most exciting up and coming improvements in Kafka Streams community and the roadmap he'd dictate if he was the benevolent dictator of Kafka Streams.

    54 min
  • Transaction Isolation - Demystified!

    If you used a relational database at all, you probably heard of transaction isolation levels. Transaction isolation levels have a massive impact on the behavior of your application - correctness, performance, and error rates. Your database may be distributed these days, so you may have to reason about distributed transactions too.

    In this video, I explain transaction isolation levels with simple examples and why these levels barely make any sense. I conclude with a few examples of the complications involved in distributed transactions - and bad news about 2 phase commits.
    ----
    For more in-depth reading:
    Transaction isolation in a few popular DBs:
    https://asktom.oracle.com/Misc/oramag...https://dev.mysql.com/doc/refman/8.0/...https://www.postgresql.org/docs/curre...
    The paper where Microsoft Research absolutely burns ANSI SQL 92 transaction isolation: https://arxiv.org/pdf/cs/0701157.pdf
    Why transactions matter: http://www.bailis.org/blog/understand...
    Super clear explanation of serialization anomalies: https://justinjaffray.com/what-does-w...
    Daniel Abadi on distributed transaction anomalies (I borrowed heavily from this in the second part of the video): https://dbmsmusings.blogspot.com/2019...
    Jepsen putting together isolation and distributed consistency: https://jepsen.io/consistency

    32 min
  • Giving and Receiving Actually Useful Advice

    YouTube and Twitter are full of “things developers should never do”. There's an endless demand for simple advice that applies in all situations.

    And that's not a bad thing. If there's a simple solution that works 80% of the time, this is valuable information. More practical than just "it depends."
    But advice-givers and advice-getters can do better. The best advice doesn't just solve an immediate problem. It tells you when to ignore this advice, and what to do if the problem isn't solved.
    I share a very old story about when I needed help sizing a connection pool, the advice I got, and the advice I now wish someone had given me.
    P.S
    Sorry about the audio quality. I had to record outside, and my lav mic wasn't as good as my desk mic.

    18 min
  • The Promise of Serverless

    When developers talk about Serverless, they often focus on FaaS. But the best Serverless experience, by far, is delivered by a data store. S3.

    Why? Because it "just works" and lets developers focus on their code.

    Serverless databases help you focus on your queries and workload. They abstract the compute. Which also means - usage based pricing.

    In this video, Ram Subramanian, Nile's CEO, joins us to discuss his vision of the perfect Serverless database experience.

    We talk about:

    - What makes S3 so amazing?
    - What would the S3 experience look like if we apply it to RDBMS?
    - Development cycle: Serverless requirements when coding, testing and finally in production.
    - Performance of Serverless databases, and what will make performance tuning a better experience
    - Elasticity and scalability. We agreed that "scale to zero" doesn't mean what everyone thinks it means.
    - Cost of Serverless. Is it actually worth it?
    - Architectures: Compute-storage separation, disaggregation, sharding and gateways.
    - Multi-regions concerns
    And of course: Do Serverless DBs make sense for SaaS developers?


    1 hr 7 min
  • Airtable - Migrating a Multitenant Architecture to MySQL 8.0

    The storage team at Airtable published a blog post describing, in detail, the migration of their petabyte-scale storage layer from MySQL 5.6 to MySQL 8.0.

    Andrew Wang, the lead of Airtable's storage team, joined us to discuss the migration, Airtable's storage architecture, data isolation levels, engineering culture, and more.

    The blog: https://medium.com/airtable-eng/migrating-airtable-to-mysql-8-0-809f0398a493

    Github's gh-ost, recommended by Andrew: https://github.com/github/gh-ost

    Andrew's LinkedIn (he's hiring): https://www.linkedin.com/in/aawang/

    39 min
  • The Multitenant journey - From 0 to 500M ARR

    SaaS applications are multi-tenant, so whether you are writing the first line of code in a new app or worried about scaling your successful SaaS fast enough - you need to be aware of multi-tenant requirements. Isolation, access control, perfornance, operations, scale and compliance   

    In this video, Ram Subramanian, Nile CEO and SaaS Community founder, shares what he learned about building multi-tenant applications, based on 20+ years of experience and 100+ customer conversations. Starting from the first decision about the data model all the way to operating large scale deployments across multiple geographies.   

    Blogs we refer to in this video: 

    https://www.notion.so/blog/sharding-postgres-at-notion 

    https://www.atlassian.com/engineering/scaling-rearchitecting-and-decomposing-confluence-cloud 

    https://www.atlassian.com/engineering/april-2022-outage-update https://slack.engineering/scaling-datastores-at-slack-with-vitess/ 

    And few other good blogs: 

    https://blog.gotenzo.com/tech/database-sharding-solving-performance-in-a-multi-tenant-restaurant-data-analytics-system

    https://blog.sentry.io/2015/07/23/transaction-id-wraparound-in-postgres/ 

    https://blog.cloudflare.com/performance-isolation-in-a-multi-tenant-database-environment/ 

    https://www.lighttag.io/blog/database-multi-tenancy/

    56 min
  • Compute-Storage Separation Explained

    Gunnar Morling asked a great querstion on Twitter: ""Separating storage and compute" vs. "Predicate push-down" -- I can't quite square these two with each other. Is there a world where they co-exist, or is it just two opposing patterns/trends in DB tech. ?"   


    Those are complimentary patterns and you definitely want them together. They appear contradictory because "separating storage and compute" can mean different things to different people.   

    In this episode, I explain the various meanings of compute-storage separation, the problems it solves, the new problem it can create, and how predicate pushdown saves the day.

    Gunnar's tweet: https://twitter.com/gunnarmorling/status/1610203324963516417
    Aurora paper: https://web.stanford.edu/class/cs245/readings/aurora.pdf

    19 min

About SaaS for Developers

From the publisher's feed

Learn what every engineer should know about building and scaling SaaS products from leaders who built world-class SaaS. We will share lessons learned, advice, tips, and great stories. This podcast is part of the SaaS community.