the bioinformatics chat

the bioinformatics chat

By Roman CheplyakaScienceLife Sciences
Download on the App Store

the bioinformatics chat episodes

  • #40 Plasmid classification and binning with Sergio Arredondo-Alonso and Anita Schürch

    Does a given bacterial gene live on a plasmid or the chromosome? What

    other genes live on the same plasmid?

    In this episode, we hear from Sergio Arredondo-Alonso and Anita Schürch, whose

    projects mlplasmids and gplas answer these types of questions.

    Links:

    • mlplasmids: a user-friendly tool to predict plasmid- and chromosome-derived sequences for single species (Sergio Arredondo-Alonso, Malbert R. C. Rogers, Johanna C. Braat, Tess D. Verschuuren, Janetta Top, Jukka Corander, Rob J. L. Willems, Anita C. Schürch)
    • gplas: a comprehensive tool for plasmid analysis using short-read graphs (Sergio Arredondo-Alonso, Martin Bootsma, Yaïr Hein, Malbert R.C. Rogers, Jukka Corander, Rob JL Willems, Anita C. Schürch)
    • 46 min
    • #39 Amplicon sequence variants and bias with Benjamin Callahan

      In this episode, Benjamin Callahan talks about some of the issues faced by

      microbiologists when conducting amplicon sequencing and metagenomic studies. The two main themes are:

      • Why one should probably avoid using OTUs (operational taxonomic units) and
      • use exact sequence variants (also called amplicon sequence variants, or
        ASVs), and how DADA2 manages to deduce the exact sequences present in the
        sample.
      • Why abundances inferred from community sequencing data are biased, and
      • how we can model and correct this bias.

        Links:

        • Exact sequence variants should replace operational taxonomic units in marker-gene data analysis (Benjamin J Callahan, Paul J McMurdie & Susan P Holmes)
        • DADA2: High-resolution sample inference from Illumina amplicon data (Benjamin J Callahan, Paul J McMurdie, Michael J Rosen, Andrew W Han, Amy Jo A Johnson & Susan P Holmes)
        • In Nature, There Is Only Diversity (Michael R. McLaren, Benjamin J. Callahan)
        • Consistent and correctable bias in metagenomic sequencing experiments (Michael R McLaren, Amy D Willis, Benjamin J Callahan)
        • 1 hr 2 min
        • #38 Issues in legacy genomes with Luke Anderson-Trocmé

          In this episode, Luke Anderson-Trocmé

          talks about his findings from the 1000 Genomes Project. Namely, the early
          sequenced genomes sometimes contain specific mutational signatures that
          haven’t been replicated from other sources and can be found via their
          association with lower base quality scores. Listen to Luke telling the story
          of how he stumbled upon and investigated these fake variants and what their
          impact is.

          Links:

          • Legacy Data Confounds Genomics Studies (bioRxiv, Molecular Biology and Evolution (paywall)) (Luke Anderson-Trocmé, Rick Farouni, Mathieu Bourgey, Yoichiro Kamatani, Koichiro Higasa, Jeong-Sun Seo, Changhoon Kim, Fumihiko Matsuda and Simon Gravel)
          • 1 hr 2 min
          • #37 Causality and potential outcomes with Irineo Cabreros

            In this episode, I talk with Irineo Cabreros about causality. We discuss why

            causality matters, what does and does not imply causality, and two
            different mathematical formalizations of causality: potential outcomes and
            directed acyclic graphs (DAGs). Causal models are
            usually considered external to and separate from statistical models, whereas
            Irineo’s new paper shows how causality can be viewed as a relationship between
            particularly chosen random variables (potential outcomes).

            Links:

            • Causal models on probability spaces (Irineo Cabreros, John D. Storey)
            • The Book of Why: The New Science of Cause and Effect (Judea Pearl, Dana Mackenzie)
            • 41 min
            • #36 scVI with Romain Lopez and Gabriel Misrachi

              In this episode, we hear from Romain Lopez and Gabriel Misrachi about

              scVI—Single-cell Variational Inference.
              scVI is a probabilistic model for single-cell gene expression data that
              combines a hierarchical Bayesian model with deep neural networks encoding the
              conditional distributions. scVI scales to over one million cells and can be
              used for scRNA-seq normalization and batch effect removal, dimensionality
              reduction, visualization, and differential expression. We also
              discuss the recently implemented in scVI automatic hyperparameter selection
              via Bayesian optimization.

              Links:

              • Deep generative modeling for single-cell transcriptomics (Romain Lopez, Jeffrey Regier, Michael Cole, Michael I. Jordan, Nir Yosef)
              • scVI on GitHub
              • Should we zero-inflate scVI?
              • Hyperparameter search for scVI
              • Droplet scRNA-seq is not zero inflated (Valentine Svensson)
              • 1 hr 21 min
              • #35 The role of the DNA shape in transcription factor binding with Hassan Samee

                Even though the double-stranded DNA has the famous regular helical shape,

                there are small variations in the geometry of the helix depending on what
                exact nucleotides its made of at that position.

                In this episode of the bioinformatics chat, Hassan Samee talks about the

                role the DNA shape plays in recognition of the DNA by DNA-binding proteins,
                such as transcription factors. Hassan also explains how his algorithm, ShapeMF,
                can deduce the DNA shape motifs from the ChIP-seq data.

                Links:

                • A De Novo Shape Motif Discovery Algorithm Reveals Preferences of Transcription Factors for DNA Shape Beyond Sequence Motifs (Md. Abul Hassan Samee, Benoit G. Bruneau, Katherine Pollard)
                • ShapeMF on GitHub
                • A picture explaining some of the DNA shape features
                • 1 hr 2 min
                • #34 Power laws and T-cell receptors with Kristina Grigaityte

                  An αβ T-cell receptor is composed of two highly variable protein chains, the α

                  chain and the β chain. However, based only on bulk DNA or RNA sequencing it is
                  impossible to determine which of the α chain and β chain sequences were paired
                  in the same receptor.

                  In this episode, Kristina Grigaityte talks about her analysis of 200,000

                  paired αβ sequences, which have been obtained by targeted single-cell RNA sequencing.
                  Kristina used the power law distribution to model the T-cell clone sizes,
                  which led her to reject the commonly held assumptions about the independence
                  of the α and β chains. We also talk about Bayesian inference of power law
                  distributions and about mixtures of power laws.

                  Links:

                  • Single-cell sequencing reveals αβ chain pairing shapes the T cell repertoire. Kristina Grigaityte, Jason A. Carter, Stephen J. Goldfless, Eric W. Jeffery, Ronald J. Hause, Yue Jiang, David Koppstein, Adrian W. Briggs, George M. Church, Francois Vigneault, Gurinder S. Atwal
                  • Bayesian inference of power law distributions. Kristina Grigaityte, Gurinder Atwal
                  • Mathematics in modern immunology. Castro M, Lythe G, Molina-París C, Ribeiro RM.
                  • Power laws, Pareto distributions and Zipf’s law. M. E. J. Newman
                  • So You Think You Have a Power Law — Well Isn’t That Special?
                  • 1 hr 27 min
                  • #33 Genome assembly from long reads and Flye with Mikhail Kolmogorov

                    Modern genome assembly projects are often based on long reads in an attempt to

                    bridge longer repeats. However, due to the higher error rate of the current
                    long read sequencers, assemblers based on de Bruijn graphs do not work well in
                    this setting, and the approaches that do work are slower.

                    In this episode, Mikhail Kolmogorov from

                    Pavel Pevzner’s lab joins us to talk about some of the ideas developed in the
                    lab that made it possible to build a de Bruijn-like assembly graph from noisy
                    reads. These ideas are now implemented in the Flye assembler, which performs
                    much faster than the existing long read assemblers without sacrificing the
                    quality of the assembly.

                    Links:

                    • Assembly of Long Error-Prone Reads Using Repeat Graphs (Mikhail Kolmogorov,
                    • Jeffrey Yuan, Yu Lin, Pavel. A. Pevzner.
                      Nature Biotechnology
                      (paywalled),
                      bioRxiv
                    • Flye on GitHub
                    • 1 hr 13 min
                    • #32 Deep tensor factorization and a pitfall for machine learning methods with Jacob Schreiber

                      In this episode, we hear from Jacob Schreiber about his algorithm,

                      Avocado.

                      Avocado uses deep tensor factorization to break a three-dimensional tensor of

                      epigenomic data into three orthogonal dimensions corresponding to cell types,
                      assay types, and genomic loci. Avocado can extract a low-dimensional,
                      information-rich latent representation from the wealth of experimental data
                      from projects like the Roadmap Epigenomics Consortium and ENCODE. This
                      representation allows you to impute genome-wide epigenomics experiments that
                      have not yet been performed.

                      Jacob also talks about a pitfall he discovered when trying to predict gene

                      expression from a mix of genomic and epigenomic data. As you increase the
                      complexity of a machine learning model, its performance may be increasing for
                      the wrong reason: instead of learning something biologically interesting, your
                      model may simply be memorizing the average gene expression for that gene
                      across your training cell types using the nucleotide sequence.

                      Links:

                      • Avocado on GitHub
                      • Multi-scale deep tensor factorization learns a latent representation of the human epigenome (Jacob Schreiber, Timothy Durham, Jeffrey Bilmes, William Stafford Noble)
                      • Completing the ENCODE3 compendium yields accurate imputations across a variety of assays and human biosamples (Jacob Schreiber, Jeffrey Bilmes, William Noble)
                      • A pitfall for machine learning methods aiming to predict across cell types (Jacob Schreiber, Ritambhara Singh, Jeffrey Bilmes, William Stafford Noble)
                      • 1 hr 16 min
                      • #31 Bioinformatics Contest 2019 with Alexey Sergushichev and Gennady Korotkevich

                        The third Bioinformatics Contest took place in

                        February 2019.

                        Alexey Sergushichev, one of the organizers of the contest,

                        and Gennady Korotkevich, the 1st prize winner,
                        join me to discuss this year’s problems.

                        Timestamps and links for the individual problems:

                        • Qualification round
                          • 00:07:14 Bee Population
                          • 00:14:12 Sequencing Errors
                          • 00:30:20 Transposable Elements
                          • Final round
                            • 00:41:35 Cancer and Chromosome Rearrangements
                            • 00:56:01 Epigenomic Marks
                            • 01:10:02 Bacterial Communities
                            • 01:27:06 Minimal Genome
                            • 01:34:56 Endangered Species
                            • Links:

                              • The contest problems on Stepik
                              • The final scoreboard
                              • Episode #18: Bioinformatics Contest 2018
                              • 1 hr 47 min

                              About the bioinformatics chat

                              From the publisher's feed

                              A podcast about computational biology, bioinformatics, and next generation sequencing.