the bioinformatics chat

the bioinformatics chat

By Roman CheplyakaScienceLife Sciences
Download on the App Store

the bioinformatics chat episodes

  • #30 Bayesian inference of chromatin structure from Hi-C data with Simeon Carstens

    Hi-C is a sequencing-based assay that provides information about the 3-dimensional organization of the genome.

    In this episode, Simeon Carstens explains how he
    applied the Inferential Structure Determination (ISD) framework to build a 3D
    model of chromatin and fit that model to Hi-C data using Hamiltonian Monte
    Carlo and Gibbs sampling.

    Links:

    • Bayesian inference of chromatin structure ensembles from population Hi-C data (Simeon Carstens, Michael Nilges, Michael Habeck)
    • Inferential Structure Determination of Chromosomes from Single-Cell Hi-C Data (Simeon Carstens, Michael Nilges, Michael Habeck)
    • 1 hr 6 min
    • #29 Haplotype-aware genotyping from long reads with Trevor Pesout

      Long read sequencing technologies, such as Oxford Nanopore and PacBio,

      produce reads from thousands to a million base pairs in length,
      at the cost of the increased error rate. Trevor Pesout
      describes how he and his colleagues leverage long reads for simultaneous
      variant calling/genotyping and phasing. This is possible thanks to a clever
      use of a hidden Markov model, and two different algorithms based on this model
      are now implemented in
      the MarginPhase and WhatsHap tools.

      Links:

      • Preprint: Haplotype-aware genotyping from noisy long reads (Jana Ebler, Marina Haukness, Trevor Pesout, Tobias Marschall, Benedict Paten)
      • 1 hr 13 min
      • #28 Space-efficient variable-order Markov models with Fabio Cunial

        This time you’ll hear from Fabio Cunial on the topic of Markov models and

        space-efficient data structures. First we recall what a Markov model is and
        why variable-order Markov models are an improvement over the standard,
        fixed-order models. Next we discuss the various data structures and indexes
        that allowed Fabio and his collaborators to represent these models in a very
        small space while still keeping the queries efficient. Burrows-Wheeler
        transform, suffix trees and arrays, tries and suffix link trees, and more!

        Links:

        • The preprint: A framework for space-efficient variable-order Markov models
        • The book: Genome-Scale Algorithm Design
        • The GitHub repo
        • 1 hr 10 min
        • #27 Classification of CRISPR-induced mutations and CRISPRpic with HoJoon Lee and Seung Woo Cho

          In this episode, HoJoon Lee and Seung Woo Cho explain how to perform a CRISPR

          experiment and how to analyze its results. HoJoon and Seung Woo developed an
          algorithm that analyzes sequenced amplicons containing the CRISPR-induced
          double-strand break site and figures out what exactly happened there (e.g.
          a deletion, insertion, substitution etc.)

          Links:

          • CRISPRpic: Fast and precise analysis for CRISPR-induced mutations via prefixed index counting
          • CRISPRpic on GitHub
          • 57 min
          • #26 Feature selection, Relief and STIR with Trang Lê

            Relief is a statistical method to perform feature selection. It could be used,

            for instance, to find genomic loci that correlate with a trait or genes whose
            expression correlate with a condition. Relief can also be made sensitive to
            interaction effects (known in genetics as epistasis).

            In this episode, Trang Lê joins me

            to talk about Relief and her version of Relief called STIR (STatistical
            Inference Relief). While traditional Relief algorithms could only rank
            features and needed a user-supplied threshold to decide which features to
            select, Trang’s reformulation of Relief allowed her to compute p-values
            and make the selection process less arbitrary.

            Links:

            • Paper: STatistical Inference Relief (STIR) feature selection
            • STIR on GitHub
            • Relief on Wikipedia
            • The original Relief paper by Kira and Rendell (1992)
            • Epistasis: what it means, what it doesn’t mean, and statistical methods to detect it in humans
            • 1 hr 9 min
            • #25 Transposons and repeats with Kaushik Panda and Keith Slotkin

              Kaushik Panda and Keith Slotkin come on the podcast to educate us about

              repetitive DNA and transposable elements. We talk LINEs, SINEs, LTRs, and even
              Sleeping Beauty transposons! Kaushik and Keith explain why repeats matter for your
              whole-genome analysis and answer listeners’ questions.

              Links:

              • Keith’s paper: The case for not masking away repetitive DNA
              • Questions for this episode on Reddit
              • 1 hr 41 min
              • #24 Read correction and Bcool with Antoine Limasset

                Antoine Limasset joins me to talk about NGS read correction.

                Antoine and his colleagues built the read correction tool Bcool based on the
                de Bruijn graph, and it corrects reads far better than any of the current methods
                like Bloocoo, Musket, and Lighter.

                We discuss why and when read correction is needed, how Bcool works, and why

                it performs better but slower than k-mer spectrum methods.

                Links:

                • Preprint: Toward perfect reads: self-correction of short reads via mapping on de Bruijn graphs
                • Bcool on GitHub
                • 1 hr
                • #23 RNA design, EteRNA and NEMO with Fernando Portela

                  In this episode, I talk to Fernando Portela,

                  a software engineer and
                  amateur scientist
                  who works on RNA design — the problem of composing an RNA sequence
                  that has a specific secondary structure.

                  We talk about how Fernando and others compete and collaborate in designing RNA

                  molecules in the online game EteRNA and about Fernando’s new
                  RNA design algorithm, NEMO, which outperforms all prior published methods by a wide margin.

                  Links:

                  • The EteRNA game
                  • The preprint about NEMO
                  • NEMO project page
                  • Single-cell RNABIO & organoids meeting in Kiev
                  • 1 hr 32 min
                  • #22 smCounter2: somatic variant calling and UMIs with Chang Xu

                    In this episode I’m joined by Chang Xu. Chang is a senior biostatistician

                    at QIAGEN and an author of smCounter2, a low-frequency somatic variant caller.
                    To distinguish rare somatic mutations from sequencing errors, smCounter2
                    relies on unique molecular identifiers, or UMIs, which help identify multiple
                    reads resulting from the same physical DNA fragment.

                    Chang explains what UMIs are, why they are useful, and how smCounter2 and other

                    tools in this space use UMIs to detect low-frequency variants.

                    Links:

                    • smCounter2 preprint
                    • smCounter2 github repository
                    • smCounter publication
                    • Review of somatic SNV callers
                    • 1 hr 5 min
                    • #21 Linear mixed models, GWAS, and lme4qtl with Andrey Ziyatdinov

                      Linear mixed models are used to analyze GWAS data and detect QTLs.

                      Andrey Ziyatdinov recently released an R package, lme4qtl, that can be used to
                      formulate and fit these models.
                      In this episode, Andrey and I discuss linear mixed models, genome-wide association studies, and strengths and weaknesses of lme4qtl.

                      Links:

                      • Paper: lme4qtl: linear mixed models with flexible covariance structure for genetic studies of related individuals
                      • lme4qtl on GitHub
                      • 51 min

                      About the bioinformatics chat

                      From the publisher's feed

                      A podcast about computational biology, bioinformatics, and next generation sequencing.