Data Engineering Podcast

MarketStore: Managing Timeseries Financial Data with Hitoshi Harada and Christopher Ryan - Episode 24


Listen Later

Summary

The data that is used in financial markets is time oriented and multidimensional, which makes it difficult to manage in either relational or timeseries databases. To make this information more manageable the team at Alapaca built a new data store specifically for retrieving and analyzing data generated by trading markets. In this episode Hitoshi Harada, the CTO of Alapaca, and Christopher Ryan, their lead software engineer, explain their motivation for building MarketStore, how it operates, and how it has helped to simplify their development workflows.

Preamble
  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
  • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
  • For complete visibility into the health of your pipeline, including deployment tracking, and powerful alerting driven by machine-learning, DataDog has got you covered. With their monitoring, metrics, and log collection agent, including extensive integrations and distributed tracing, you’ll have everything you need to find and fix performance bottlenecks in no time. Go to dataengineeringpodcast.com/datadog today to start your free 14 day trial and get a sweet new T-Shirt.
  • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
  • Your host is Tobias Macey and today I’m interviewing Christopher Ryan and Hitoshi Harada about MarketStore, a storage server for large volumes of financial timeseries data
  • Interview
    • Introduction
    • How did you get involved in the area of data management?
    • What was your motivation for creating MarketStore?
    • What are the characteristics of financial time series data that make it challenging to manage?
    • What are some of the workflows that MarketStore is used for at Alpaca and how were they managed before it was available?
    • With MarketStore’s data coming from multiple third party services, how are you managing to keep the DB up-to-date and in sync with those services?
      • What is the worst case scenario if there is a total failure in the data store?
      • What guards have you built to prevent such a situation from occurring?

      • Since MarketStore is used for querying and analyzing data having to do with financial markets and there are potentially large quantities of money being staked on the results of that analysis, how do you ensure that the operations being performed in MarketStore are accurate and repeatable?

      • What were the most challenging aspects of building MarketStore and integrating it into the rest of your systems?

      • Motivation for open sourcing the code?

      • What is the next planned major feature for MarketStore, and what use-case is it aiming to support?

      • Contact Info
        • Christopher
          • Email

          • Hitoshi

            • Email

            • Parting Question
              • From your perspective, what is the biggest gap in the tooling or technology for data management today?
              • Links
                • MarketStore
                  • GitHub
                  • Release Announcement

                  • Alpaca

                  • IBM

                  • DB2

                  • GreenPlum

                  • Algorithmic Trading

                  • Backtesting

                  • OHLC (Open-High-Low-Close)

                  • HDF5

                  • Golang

                  • C++

                  • Timeseries Database List

                  • InfluxDB

                  • JSONRPC

                  • Slait

                  • CircleCI

                  • GDAX

                  • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                    Support Data Engineering Podcast

                    ...more
                    View all episodesView all episodes
                    Download on the App Store

                    Data Engineering PodcastBy Tobias Macey

                    • 4.5
                    • 4.5
                    • 4.5
                    • 4.5
                    • 4.5

                    4.5

                    142 ratings


                    More shows like Data Engineering Podcast

                    View all
                    This Week in Startups by Jason Calacanis

                    This Week in Startups

                    1,297 Listeners

                    The Changelog: Software Development, Open Source by Changelog Media

                    The Changelog: Software Development, Open Source

                    288 Listeners

                    The a16z Show by Andreessen Horowitz

                    The a16z Show

                    1,109 Listeners

                    Software Engineering Daily by Software Engineering Daily

                    Software Engineering Daily

                    630 Listeners

                    Risky Business by Risky Business Media

                    Risky Business

                    372 Listeners

                    Talk Python To Me by Michael Kennedy

                    Talk Python To Me

                    583 Listeners

                    Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                    Super Data Science: ML & AI Podcast with Jon Krohn

                    308 Listeners

                    NVIDIA AI Podcast by NVIDIA

                    NVIDIA AI Podcast

                    345 Listeners

                    Syntax - Tasty Web Development Treats by Wes Bos & Scott Tolinski - Full Stack JavaScript Web Developers

                    Syntax - Tasty Web Development Treats

                    986 Listeners

                    Practical AI by Practical AI LLC

                    Practical AI

                    207 Listeners

                    Dwarkesh Podcast by Dwarkesh Patel

                    Dwarkesh Podcast

                    552 Listeners

                    The Data Engineering Show by The Firebolt Data Bros

                    The Data Engineering Show

                    10 Listeners

                    Latent Space: The AI Engineer Podcast by Latent.Space

                    Latent Space: The AI Engineer Podcast

                    103 Listeners

                    This Day in AI Podcast by Michael Sharkey, Chris Sharkey

                    This Day in AI Podcast

                    229 Listeners

                    The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

                    The AI Daily Brief: Artificial Intelligence News and Analysis

                    686 Listeners