the bioinformatics chat

the bioinformatics chat

By Roman CheplyakaScienceLife Sciences
Download on the App Store

the bioinformatics chat episodes

  • #20 B cell receptor substitution profile prediction and SPURF with Kristian Davidsen and Amrit Dhar

    In this episode

    Kristian Davidsen
    and
    Amrit Dhar
    present their project called SPURF.
    SPURF can predict the B cell receptor (BCR) substitution profile of a given clonal
    family based on a single representative sequence from that family.
    SPURF works by fitting a tensor regression model to publicly available
    Rep-seq data.

    Links:

    • Preprint: Predicting B Cell Receptor Substitution Profiles Using Public Repertoire Data
    • Blog post about SPURF by Erick Matsen
    • SPURF on GitHub
    • 2 hr 2 min
    • #19 Genome fingerprints with Gustavo Glusman

      In this episode, Gustavo Glusman explains his method of reducing a VCF file

      to a small “fingerprint”, which could be then used to detect duplicate genomes,
      infer relatedness, map the population structure, and more.

      Links:

      • The genome fingerprints paper
      • The genotype fingerprints preprint
      • The data fingerprints preprint
      • The blog post about time series visualization
      • 1 hr 29 min
      • #18 Bioinformatics Contest 2018 with Alexey Sergushichev and Ekaterina Vyahhi

        The final round of Bioinformatics Contest 2018 was held on February 24-25th,

        and the qualification round took place two weeks earlier.

        I invited the organizers of the contest, Alexey Sergushichev and Ekaterina Vyahhi,

        to discuss the problems and find out what it was like to organize the contest.

        Timestamps for the problems:

        • Qualification round
          • 0:41:38 Problem 1. Synthesis of ATP
          • 0:48:46 Problem 2. Restriction Sites
          • 1:06:42 Problem 3. Tandem Repeats
          • Final round
            • 1:14:00 Problem 1. Recombination of Plasmids
            • 1:25:30 Problem 2. Species Recovering
            • 1:28:39 Problem 3. Haplotype Phasing
            • 1:36:34 Problem 4. Cluster the Reads
            • 1:40:25 Problem 5. Cattle Breeding
            • Links:

              • The contest problems on Stepik
              • The final scoreboard
              • Bioinformatics Institute
              • 1 hr 54 min
              • #17 Rarefaction, alpha diversity, and statistics with Amy Willis

                In this episode, Amy Willis joins me to talk about good and bad ways to

                estimate taxonomic richness in microbial ecology studies.

                Links:

                • Rarefaction, alpha diversity, and statistics
                • Estimating Diversity via Frequency Ratios
                • Estimating the Number of Species in Microbial Diversity Studies
                • Summer Institutes 2018 at the University of Washington
                • STAMPS: Strategies and Techniques for Analyzing Microbial Population Structures
                • bio2040, a new podcast by Flavio Rump
                • 1 hr 15 min
                • #16 Javier Quilez on what makes large sequencing projects successful

                  Javier Quilez and I discuss what it’s like to be

                  a bioinformatician, how to improve communication between the wet and dry labs
                  and make the research more reproducible.

                  Make sure to read Javier’s paper we are discussing; it’s a light and entertaining read.

                  The last author on this paper is Guillaume Filion, whom you may
                  remember from the episode on generating functions.

                  Links:

                  • Parallel sequencing lives, or what makes large sequencing projects successful
                  • 1 hr 4 min
                  • #14 Generating functions for read mapping with Guillaume Filion

                    Guillaume Filion recently published a preprint in which he applies

                    generating functions, a concept from analytic combinatorics,
                    to estimating the optimal seed length for read mapping.

                    In this episode, Guillaume and I attempt to explain the core concepts from

                    analytic combinatorics and why they are useful in modeling sequences.

                    Links:

                    • Guillaume’s preprint: Analytic combinatorics for bioinformatics I: seeding methods
                    • Once upon a BLAST
                    • Guillaume’s blog, «The Grand Locus»
                    • Dan Gusfield’s home page
                    • featuring the fast fourier transform lectures I mention in the podcast

                      After we recorded the podcast, Guillaume wrote to me to clarify the

                      relationship between read mapping and BLAST:

                      I looked into my notes about BLAST. The problem that it solves is the following: “Given that a local alignment has score S, what is the probability that it does not contain a word of score T or greater”? The background work of Karlin and Altschul is used to give a statistical significance for S (what is the probability that a “Smith-Waterman random walk” starting at height 0 would reach height S, i.e. what is the probability that aligning two random proteins would yield a score S). The authors write in the original paper “Theory does not yet exist to calculate the probability q that such segment pair will contain a word pair with a score of at least T. However, one argument suggests that q should depend exponentially upon the score of the MSP”.

                      This is the part that I did not remember well. MSP stands for Maximal Segment Pair, this is the “longest fragment” with “highest score” in the alignment. I thought that Karlin and Altschul solved this part as well, but the authors just go empirical and they calibrate the relationship between T and S with simulations.

                      I realize a little bit better now that my work is precisely about this problem that the authors of BLAST could not solve, but as you pointed out, I am attacking only a very specific sub-case that is much easier because the models of sequencing error are much simpler than protein evolution. BLAST is concerned with local alignment, so it wants to get all the hits with an MSP score above S. Short read mapping just wants the true location of the read, which does not really have the notion of a score S. But still, mathematically, it is equivalent to the case where S is a constant that depends only on the read size and the distribution of the score T depends only on the seed length and the error rate. I have a few ideas of how to use analytic combinatorics to solve the problem for proteins, but it is mostly complicated because the variable of interest T is a fractional numbers and not an integer…

                      So what is different from BLAST? The right answer (I think) is that BLAST finds all the hits with an MSP above statistical background, but it says nothing of the probability that the true location contains such an MSP, so it is hard to calibrate the heuristic for that specific problem. In reality, the parallel with BLAST is just the basic strategy: make a statistical model for your problem and use it to calibrate the heuristic.

                      1 hr 11 min
                    • #13 Bracken with Jennifer Lu

                      Jennifer Lu joins me to discuss species abundance estimation from metagenomic sequencing data.

                      Links:

                      • The Bracken paper
                      • The Kraken paper
                      • The preprint that applies Kallisto to metagenomics
                      • 47 min
                      • #12 Modelling the immune system and C-ImmSim with Filippo Castiglione

                        In this episode, Filippo Castiglione and I discuss different ways to model the immune system.

                        Links:

                        • Celada’s and Seiden’s 1992 paper, “A computer model of cellular interactions in the immune system”
                        • Filippo’s 2001 paper, “Design and implementation of an immune system simulator”
                        • PLOS One paper, “Computational Immunology Meets Bioinformatics: The Use of Prediction Tools for Molecular Binding in the Simulation of the Immune System”
                        • A paper about modelling sea bass vaccination using C-ImmSim
                        • C-ImmSim homepage
                        • The online simulator based on C-ImmSim and the paper describing it
                        • Special thanks to Martina Stoycheva for bringing this work to my attention.

                          1 hr 7 min
                        • #11 Collective cell migration with Linus Schumacher

                          In this episode, Linus Schumacher joins me to discuss mathematical models of collective cell migration and multidisciplinary research.

                          Links:

                          • Semblance of Heterogeneity in Collective Cell Migration and associated code
                          • Multidisciplinary approaches to understanding collective cell migration in developmental biology
                          • Linus’s homepage
                          • 1 hr 1 min

                          About the bioinformatics chat

                          From the publisher's feed

                          A podcast about computational biology, bioinformatics, and next generation sequencing.