Micro binfie podcast

Micro binfie podcast

By Microbial BioinformaticsScience
Download on the App Store

Micro binfie podcast episodes

  • 116 AI Authorship and Ethics in Academic Publishing for Genomics
    The podcast discusses an article co-authored by Andrew Page, examining the use of GPT-4 for research publication. The conversation focuses on the authorship of articles generated by GPT-4 and the implications for academic publishing.
    Authorship and Ethics:
    Andrew discusses the question of authorship when AI-generated content is involved in research articles. He explores the ethical implications and potential biases associated with AI-assisted writing, such as the omission of minority figures and novel discoveries. He emphasizes the importance of transparency when using AI and its potential to democratize research, as long as ethical guidelines are maintained.
    AI & Scientific Journals:
    The podcast delves into the current landscape of AI in academic publishing. It addresses the commercial use of AI in crafting manuscripts for research articles and the necessity of distinguishing between manual and AI-generated contributions. The possible misalignment of GPT-4's commercial objectives with academic goals is highlighted.
    Risks and Benefits:
    Andrew outlines the risks of using AI in publishing, such as unintentional plagiarism, biases, and outdated methods. He provides an example of bioinformatics software recommending deprecated methods, illustrating the need for caution. The conversation also touches upon the AI's potential to introduce bias unintentionally, citing past incidents where AI models quickly adopted extremist views.
    Andrew's co-authors, Niamh Tumelty and Sam Sheppard, bring different perspectives on ethics and the impact of AI on publishing. Niamh, associated with the London School of Economics, emphasizes ethical considerations, while Sam, editor-in-chief of Microbial Genomics, underscores the need to adapt to the reality of AI contributions in journal submissions.
    In conclusion, the podcast underscores the importance of recognizing and navigating the ethical challenges posed by AI in academic publishing. It suggests that the technology may evolve faster than policies can adapt, necessitating an ongoing conversation among researchers, publishers, and AI developers.
    Links:
    https://microbiologysociety.org/blog/microbe-talk-ai-a-useful-tool-or-dangerous-unstoppable-force.html
    https://www.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.001049
    33 min
  • 115 Write-the: speeding up software development for bioinformatics
    We continue our conversation with Wytamma Wirth about write-the and all things AI. It starts with discussing the usage of language models, specifically ChatGPT, in writing boilerplate code, and how it can assist in generating code snippets, unit tests, and even documentation strings. The participants also explore the potential of incorporating it into code editors to make coding more efficient and less error-prone.
    The conversation then shifts to discuss the generation of research papers, specifically software announcements, by leveraging code documentation. The participants believe ChatGPT could be useful in generating introductions and backgrounds for such publications. They also touch upon the utility of language models in translating documentation into different human languages to assist non-native English speakers.
    The discussion returns to code documentation, focusing on the tool "write the docs" which auto-generates well-structured and searchable documentation websites. The participants appreciate the tool's ease of use and the potential it has in maintaining proper documentation for projects. The conversation ends with an acknowledgment of the importance of human oversight in automating tasks using language models.
    Links:
    Write-the software: https://github.com/Wytamma/write-the
    Wytamma Wirth: https://www.wytamma.com/
    25 min
  • 114 Write-the: Automating Code Documentation ChatGPT
    In this episode, we dive deep into the world of automated code documentation and conversion using ChatGPT through the write-the software developed by Dr Wytamma Wirth from The University of Melbourne. Our guest, an experienced software engineer, takes us on a journey through the challenges and nuances of writing code documentation and the role AI can play in easing this process. We explore the intersection of ChatGPT's capabilities with Write the Docs, a documentation system widely used by developers. From highlighting ChatGPT's ability to understand and generate code snippets, to demonstrating real-time code conversion across multiple programming languages, this episode is a treasure trove for developers looking to enhance their workflow. Whether you're a seasoned developer or just getting started, tune in to discover how the synergy of AI and coding can elevate your documentation game to the next level!
    Links:
    Write-the software: https://github.com/Wytamma/write-the
    Wytamma Wirth: https://www.wytamma.com/
    26 min
  • 110 ChatGPT: The Bioinformatics Calculator
    In this episode there is a comprehensive discussion on the influence of AI, especially GPT-4, in the sphere of microbial bioinformatics. They reflect on a study testing GPT-4's problem-solving capabilities, which raises concerns about its potential impact on employment practices and academic integrity.
    There's speculation that AI's proficiency in tackling standard technical problems could interfere with genuinely evaluating a candidate's knowledge during interviews. Drawing parallels with calculators, the hosts deliberate on whether AI tools should be permitted during assessments. They stress the necessity for individuals to possess a deep understanding of their domain to accurately interpret and validate AI solutions.
    Discussing the AI's limitations, the hosts highlight its struggles with regular expressions and handling larger scripts. They observe the AI tends to loop and repeat itself, performing better with shorter scripts but faltering on more complex tasks often seen in bioinformatics. This prompts a discussion on how educators should address these developments in their teaching strategies.
    Moreover, the hosts explore the potential of large language models to improve base calling and read correction in sequencing, drawing on the structured and predictable nature of language and genetic code. They also discuss the idea of introducing randomness in these models to generate creative and varied solutions, potentially predicting future alleles or gene configurations.
    Ultimately, they express a blend of enthusiasm and apprehension towards the swift advances in this field and the ensuing implications for bioinformatics. They end on a note of anticipation for future developments, with a humorous nod towards AI's potential for automating mundane tasks like auto-correcting sample sheets.
    References:
    What Is ChatGPT Doing … and Why Does It Work? https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-doing-and-why-does-it-work/
    Many bioinformatics programming tasks can be automated with ChatGPT
    https://arxiv.org/ftp/arxiv/papers/2303/2303.13528.pdf
    ChatGPT for bioinformatics
    https://medium.com/@91mattmoore/chatgpt-for-bioinformatics-404c6d0817a1
    Empowering Beginners in Bioinformatics with ChatGPT https://www.biorxiv.org/content/10.1101/2023.03.07.531414v1
    Lawyer uses GPT and get ethics violation https://simonwillison.net/2023/May/27/lawyer-chatgpt/
    Can ChatGPT solve bioinformatic problems with Python?
    https://dmnfarrell.github.io/bioinformatics/chatGPT-python
    27 min
  • 109 AI Unleashed: Navigating the Opportunities and Challenges of AI in Microbial Bioinformatics
    In this episode of the Micro Binfie Podcast, titled "AI Unleashed: Navigating the Opportunities and Challenges of AI in Microbial Bioinformatics", Lee, Nabil, and Andrew unpack the implications of generative predictive text AI tools, notably GPT, on microbial bioinformatics.
    They kick off the conversation by outlining the various applications of AI tools in their work, which range from generating boilerplate programs, drafting documents, to summarizing vast tracts of data. Andrew talks about his experience with GPT in coding, specifically via VS Code and GitHub Copilot, highlighting how GPT can generate nearly 90% of the necessary code based on a brief description of the task, thereby accelerating his work.
    He goes on to discuss the use of GPT in clarifying lines of code and notes that they used AI to generate a paper on the ethical considerations of employing AI in microbial genomics research during a recent hackathon. The conversation then switches gears as Nabil shares his experience of using GPT to standardize date formats in tables and summarize paper abstracts. While GPT is generally accurate in performing simple tasks, he warns that the tool can sometimes provide erroneous answers.
    Nabil also highlights GPT's ability to generate plausible but inaccurate responses for complex prompts, as illustrated by his experience when he used it to find a route in a video game. Andrew then talks about a script they created during a hackathon, which produces podcast episodes reviewing math tools. He points out the issues encountered, such as GPT providing wrong factual information.
    Looking ahead, Andrew envisions a future awash with GPT-generated content that may or may not be correct, raising the challenge of discerning real and false information. However, they also acknowledge the potential benefits of AI technologies for those with visual impairments, though it's far from a perfect solution at present.
    The conversation veers to the use of AI tech in handling boilerplate code and generating code snippets based on predictive text. The hosts further discuss the potential for this tool in rapid language learning. A live experiment ensues where Nabil and Andrew use a Perl script and utilize GPT-4 to convert this script into Python and back again to assess its capabilities in language translation. The AI tool proves proficient, considering comments, usage, and authorship and employing popular libraries like BioPython intelligently, though it does leave a disclaimer about potential inaccuracies.
    They consider the possibility of using AI to optimize coding, similar to minifying JavaScript, and even the idea of iterating through multiple languages and assessing the output. Nabil initiates a simpler task for the AI, asking it to write a Python script translating DNA into protein, which then gets translated into Rust. Andrew shares his experience of using AI to generate a Python class that compares two spreadsheets using pandas, demonstrating AI's comprehension and execution of complex tasks.
    In summary, this episode underscores the power and potential of AI in coding and the need for human oversight to ensure the quality and effectiveness of AI-generated content. It offers a glimpse into a future where AI tools, despite their limitations, can revolutionize many aspects of programming, bringing in new efficiencies and methods of working.
    31 min
  • Encore: What language should I learn?
    The MicroBinfie podcast discusses the top programming languages for bioinformatics. Andrew, Lee, and Nabil agree that Python is a great starting point for its consistency and rigor. Its strict syntax is ideal for teaching programming fundamentals that are essential in any language. In contrast, Perl encourages multiple ways of doing the same thing, creating confusion and difficulties in keeping track of things.
    The hosts caution against starting with trendy languages that are constantly changing. Instead, stick with more established languages like Python, which have established libraries and concepts that will help you advance more easily. Trendy languages come and go like changing tides, making them riskier choices. Additionally, they highlight the importance of understanding databases and their primary keys and unique fields. SQL is useful, particularly in dealing with large datasets. It is consistent across flavors and unlikely to go away soon. It takes a lot of skill to optimize queries to work in milliseconds.
    The hosts emphasize that the language you choose to learn depends on your individual goals and environment. For instance, Lee suggests that you should look to who is in your space and what they are using and who is willing to help you. Once you understand the programming concepts, it is easier to transfer them to other languages, and it is just a question of understanding the syntax.
    Andrew, Lee, and Nabil also discuss their own trajectories of learning programming languages, revealing that it takes a long time to become an expert in a language, and it is something that needs to be appreciated. They highlight the difference between just learning the basics of a language and really getting into the depths of it and the frameworks and libraries.
    The hosts also mention languages that are important to pick up, like SQL and bash scripting, and languages that are popular for web development, like JavaScript. However, they caution that JavaScript and Java are not the same thing and that JavaScript has a reputation for being a weird language.
    When asked what language they would choose for a task, Nabil says he would use Perl, Lee mentions R for stats, while Andrew admits that he has to relearn R every time he comes back to it and therefore prefers Perl for quick scripts. They also discuss their love-hate relationship with R, mentioning that while it has useful libraries like GGplot and GGtree, its syntax is difficult to work with and has separate paradigms of approaching the same problem.
    The hosts conclude by acknowledging that there is no one-size-fits-all approach to learning programming languages. One should choose based on their goals, environment, and personal preferences. Python is a useful language to learn, even if one is not interested in bioinformatics. Additionally, they note that the fundamentals of databases and how they work are crucial to understand and utilized across fields.
    30 min
  • 108 SeqCode: a nomenclatural code for prokaryotes described from sequence data
    We are back talking about systematics, and SeqCode; a nomenclatural code for prokaryotes described from sequence data.
    Marike Palmer is a Postdoctoral researcher in the School of Life Sciences at the University of Nevada Las Vegas and Miguel Rodriguez is an Assistant Professor of Bioinformatics at the University of Innsbruck in the departments of Microbiology and the Digital Science Center (DiSC).
    Link to paper: https://www.nature.com/articles/s41564-022-01214-9
    History paper: https://www.sciencedirect.com/science/article/pii/S0723202022000121
    They discussed the SeqCode, a nomenclature code for Prokaryotes described from sequence data. The SeqCode was created to provide a specific nomenclature code for previously uncultivated organisms. Palmer explained that the impetus for the SeqCode was the need to accommodate previously uncultivated organisms under a specific nomenclature code. She emphasized that the SeqCode was written to allow any peer-reviewed publication, but noted that the authors have designed three paths of validation in the SeqCode. They hope that anyone proposing a name will work with the curriculum team to ensure the best quality descriptions, names, etymology, and solidification.
    Rodriguez discussed the SeqCode's governance, which is already in place, and they have made them public so that anyone interested can join the SeqCode community. The governance structure comprises an executive board, committees, and working groups. The position's co-opted members hold some of the committees of these committees, while some are chosen by ballot.
    The hosts sought to clarify the relationship between the Isme Society, which is backing the SeqCode, and the wider field in general. Rodriguez explained that ISME is simply providing support as an umbrella organization for the SeqCode.
    Palmer and Rodriguez clarified that the SeqCode is not a competing code but rather a parallel one that aims to accommodate previously uncultivated organisms. The SeqCode was created to provide a specific nomenclature code for previously uncultivated organisms. Palmer noted that most scientists culture prokaryotes not for naming but to advance their knowledge of these organisms through physiology experiments. They emphasized that the new system is the result of a long collaborative effort that involved many different viewpoints and philosophies.
    The episode also discussed the practical requirements for naming under the new system, which include standards for the completeness and contamination levels required in the genome sequence data. Palmer noted that while the 16S rRNA gene sequence was not required for naming, it was recommended for improved accuracy in cross-talk between different taxonomies. The conversation highlighted the importance and challenges of naming microorganisms and the ongoing efforts to create a system that is inclusive of all microorganisms, both cultivated and uncultivated.
    Rodriguez and Palmer also discussed the SeqCode, a nature code for naming prokaryotes described from sequence data. They agreed that high-quality genomes should be the main control types to ensure the system builds up rather than breaks down. They noted the challenge of obtaining full genomes of some organisms, such as obligate intracellular parasites but suggested obtaining housekeeping genes as a potential solution. They further explained the technical issue of estimating completeness or contamination for many taxa, but Palmer confirmed that registering a name on the SeqCode registry requires adding such estimates.
    It emphasized the importance of collaboration within the scientific community and the need to create a system that is inclusive of all microorganisms. It also highlighted the challenges inherent in the process of naming microorganisms but demonstrated that it is an ongoing process, and that scientists are working to create a system that is accurate, practical, and beneficial for all.
    46 min

About Micro binfie podcast

From the publisher's feed

Microbial Bioinformatics is a rapidly changing field marrying computer science and microbiology. Join us as we share some tips and tricks we’ve learnt over the years. If you’re student just getting to grips to the field, or someone who just wants to keep tabs on the latest and greatest - this podcast is for you.

More shows like Micro binfie podcast

Unexplainable by Vox

Unexplainable

2,318 Listeners

The Rest Is Politics by Goalhanger

The Rest Is Politics

3,024 Listeners

The Rest Is Entertainment by Goalhanger

The Rest Is Entertainment

785 Listeners