the bioinformatics chat

the bioinformatics chat

By Roman CheplyakaScienceLife Sciences
Download on the App Store

the bioinformatics chat episodes

  • #50 ENCODE3 with Jill Moore

    In this episode, Jacob Schreiber interviews Jill Moore about

    recent research from the ENCODE Project. They begin their
    discussion with an overview and goals of the ENCODE Project, and then
    discuss a bundle of papers that were recently published in various
    Nature journals and the flagship paper, Expanded encyclopaedias of DNA elements in the human and mouse genomes.
    They conclude their discussion by talking about the challenges with
    managing a large project as a trainee in a consortium setting.

    Links:

    • Expanded encyclopaedias of DNA elements in the human and mouse genomes (The ENCODE Project Consortium, Jill E. Moore, […], Zhiping Weng)
    • SCREEN
    • The ENCODE Portal
    • The ENCODE3 Publication Bundle
    • 57 min
    • #49 Most Permissive Boolean Networks with Loïc Paulevé

      In systems biology, Boolean networks are a way to model interactions such as

      gene regulation or cell signaling. The standard
      interpretations of Boolean networks are the synchronous, asynchronous, and
      fully asynchronous semantics.

      In this episode, Loïc Paulevé explains how the

      same Boolean networks can be interpreted in a new, “most permissive” way.
      Loïc proved mathematically that his semantics can reproduce all behaviors
      achievable by a compatible quantitative model, whereas the
      traditional interpretations in general cannot. Furthermore, it turns out that
      deciding whether a certain state in a Boolean network is reachable can be done
      much more efficiently in MPBNs than in the traditional interpretations.

      Links:

      • Reconciling Qualitative, Abstract, and Scalable Modeling of Biological Networks (Loïc Paulevé, Juraj Kolčák, Thomas Chatain, Stefan Haar)
      • mpbn on GitHub: an implementation of reachability and attractor analysis in Most Permissive Boolean Networks
      • BoNesis on GitHub: synthesis of Most Permissive Boolean Networks from network architecture and dynamical properties
      • 1 hr 5 min
      • #48 Machine learning for drug development with Marinka Zitnik

        In this episode, Jacob Schreiber interviews Marinka Zitnik about

        applications of machine learning to drug development.
        They begin their discussion with an overview of open research questions in the
        field, including limiting the search space of high-throughput testing methods,
        designing drugs entirely from scratch, predicting ways that existing drugs can
        be repurposed, and identifying likely side-effects of combining existing drugs
        in novel ways. Focusing on the last of these areas, they then discuss one of
        Marinka’s recent papers, Modeling polypharmacy side effects with graph
        convolutional networks.

        Links:

        • Modeling polypharmacy side effects with graph convolutional networks (Marinka Zitnik, Monica Agrawal, Jure Leskovec)
        • Network Medicine Framework for Identifying Drug Repurposing Opportunities for COVID-19 (Deisy Morselli Gysi, Ítalo Do Valle, Marinka Zitnik, Asher Ameli, Xiao Gan, Onur Varol, Helia Sanchez, Rebecca Marlene Baron, Dina Ghiassian, Joseph Loscalzo, Albert-László Barabási)
        • AI Cures initiative
        • 1 hr 26 min
        • #47 Reproducible pipelines and NGLess with Luis Pedro Coelho

          NGLess is a programming language specifically

          targeted at next generation sequencing (NGS) data processing.
          In this episode we chat with its main developer, Luis Pedro
          Coelho, about the benefits of domain-specific
          languages, pros and cons of Haskell in bioinformatics, reproducibility, and of
          course NGLess itself.

          Links:

          • NGLess on GitHub
          • NG-meta-profiler: fast processing of metagenomes using NGLess, a
          • domain-specific language
            (Luis Pedro Coelho, Renato Alves, Paulo Monteiro, Jaime Huerta-Cepas, Ana Teresa Freitas, Peer Bork)
            58 min
          • #46 HiFi reads and HiCanu with Sergey Nurk and Sergey Koren

            In this episode, I continue to talk (but mostly listen) to Sergey Koren and Sergey Nurk.

            If you missed the previous episode, you should probably start there.
            Otherwise, join us to learn about HiFi reads, the tradeoff between read length
            and quality, and what tricks HiCanu employs to resolve highly similar repeats.

            Links:

            • HiCanu: accurate assembly of segmental duplications, satellites, and allelic variants from high-fidelity long reads (Sergey Nurk, Brian P. Walenz, Arang Rhie, Mitchell R. Vollger, Glennis A. Logsdon, Robert Grothe, Karen H. Miga, Evan E. Eichler, Adam M. Phillippy, Sergey Koren)
            • Canu on GitHub (includes the HiCanu mode)
            • The Telomere-to-Telomere (T2T) consortium
            • 1 hr 10 min
            • #45 Genome assembly and Canu with Sergey Koren and Sergey Nurk

              In this episode, Sergey Nurk and Sergey Koren from the NIH share their thoughts

              on genome assembly. The two Sergeys tell the stories behind their amazing
              careers as well as behind some of the best known genome assemblers: Celera
              assembler, Canu, and SPAdes.

              Links:

              • Canu on GitHub
              • SPAdes on GitHub
              • 1 hr 17 min
              • #44 DNA tagging and Porcupine with Kathryn Doroschak

                Porcupine is a molecular tagging system—a way to tag physical

                objects with pieces of DNA called molecular bits, or molbits for short.
                These DNA tags then can be rapidly sequenced on an Oxford Nanopore MinION
                device without any need for library preparation.

                In this episode, Katie Doroschak explains how Porcupine works—how molbits

                are designed and prepared, and how they are directly recognized by the
                software without an intermediate basecalling step.

                Links:

                • Porcupine: Rapid and robust tagging of physical objects using nanopore-orthogonal DNA strands (Kathryn Doroschak, Karen Zhang, Melissa Queen, Aishwarya Mandyam, Karin Strauss, Luis Ceze, Jeff Nivala)
                • 45 min
                • #43 Generalized PCA for single-cell data with William Townes

                  Will Townes proposes a new, simpler way to analyze scRNA-seq data with unique

                  molecular identifiers (UMIs). Observing that such data is not zero-inflated,
                  Will has designed a PCA-like procedure inspired by generalized linear models
                  (GLMs) that, unlike the standard PCA, takes into account statistical
                  properties of the data and avoids spurious correlations (such as one or more
                  of the top principal components being correlated with the number of non-zero
                  gene counts).

                  Also check out Will’s paper for a feature selection algorithm based on

                  deviance, which we didn’t get a chance to discuss on the podcast.

                  Links:

                  • Feature selection and dimension reduction for single-cell RNA-Seq based on a multinomial model (F. William Townes, Stephanie C. Hicks, Martin J. Aryee, Rafael A. Irizarry)
                  • GLM-PCA for R
                  • GLM-PCA for Python
                  • scry: an R package for feature selection by deviance (alternative to highly variable genes)
                  • Droplet scRNA-seq is not zero-inflated (Valentine Svensson)
                  • 1 hr
                  • #42 Spectrum-preserving string sets and simplitigs with Amatur Rahman and Karel Břinda

                    In this episode, we hear from Amatur Rahman

                    and Karel Břinda, who
                    independently of one another released preprints on the same concept, called
                    simplitigs or spectrum-preserving string sets. Simplitigs offer a way to
                    efficiently store and query large sets of k-mers—or, equivalently, large de
                    Bruijn graphs.

                    Links:

                    • Simplitigs as an efficient and scalable representation of de Bruijn graphs (Karel Břinda, Michael Baym, Gregory Kucherov)
                    • Representation of k-mer sets using spectrum-preserving string sets (Amatur Rahman, Paul Medvedev)
                    • Open mic
                    • 54 min
                    • #41 Epidemic models with Kris Parag

                      Kris Parag is here to teach us about the mathematical modeling of

                      infectious disease epidemics. We discuss the SIR model, the renewal models, and how
                      insights from information theory can help us predict where an epidemic is
                      going.

                      Links:

                      • Optimising Renewal Models for Real-Time Epidemic Prediction and Estimation (KV Parag, CA Donnelly)
                      • Adaptive Estimation for Epidemic Renewal and Phylogenetic Skyline Models (KV Parag, CA Donnelly)
                      • The listener survey
                      • 1 hr 9 min

                      About the bioinformatics chat

                      From the publisher's feed

                      A podcast about computational biology, bioinformatics, and next generation sequencing.