NLP Highlights

NLP Highlights

By Allen Institute for Artificial IntelligenceScience
Download on the App Store

NLP Highlights episodes

  • 84 - Large Teams Develop, Small Groups Disrupt, with Lingfei Wu
    In a recent Nature paper, Lingfei Wu (Ling) suggests that smaller teams of scientists tend to do more disruptive work. In this episode, we invite Ling to discuss their results, how they define disruption and possible reasons why smaller teams may be better positioned to do disruptive work. We also touch on robustness of the disruption metric, differences between research disciplines, and sleeping beauties in science.
    Lingfei Wu’s homepage: https://www.knowledgelab.org/people/detail/lingfei_wu/
    Paper: https://www.nature.com/articles/s41586-019-0941-9
    Note: Lingfei is on the job market for faculty positions at the intersection of social science, computer science and communication.
    39 min
  • 83 - Knowledge Base Construction, with Sebastian Riedel
    In this episode, we invite Sebastian Riedel to talk about knowledge base construction (KBC). Why is it an important research area? What are the tradeoffs between using an open vs. closed schema? What are popular methods currently used, and what challenges prevent the adoption of KBC methods? We also briefly discuss the AKBC workshop and its graduation into a conference in 2019.
    Sebastian Riedel's homepage: http://www.riedelcastro.org/
    AKBC conference: http://www.akbc.ws/2019/
    39 min
  • 82 - Visual Reasoning, with Yoav Artzi
    In this episode, Yoav Artzi joins us to talk about visual reasoning. We start by defining what visual reasoning is, then discuss the pros and cons of different tasks and datasets. We discuss some of the models used for visual reasoning and how they perform, before ending with open questions in this young, exciting research area.
    Yoav Artzi: https://yoavartzi.com/
    NLVR: https://github.com/clic-lab/nlvr/tree/master/nlvr
    NLVR2: https://github.com/clic-lab/nlvr/tree/master/nlvr2
    CLEVR dataset: https://cs.stanford.edu/people/jcjohns/clevr/
    VQA: https://visualqa.org/
    GQA: https://cs.stanford.edu/people/dorarad/gqa/index.html
    Neural module networks: https://arxiv.org/abs/1511.02799
    43 min
  • 81 - BlackboxNLP, with Afra Alishahi and Tal Linzen
    Neural models recently resulted in large performance improvements in various NLP problems, but our understanding of what and how the models learn remains fairly limited. In this episode, Tal Linzen and Afra Alishahi talk to us about BlackboxNLP, an EMNLP’18 workshop dedicated to the analysis and interpretation of neural networks for NLP. In the workshop, computer scientists and cognitive scientists joined forces to probe and analyze neural NLP models.
    BlackboxNLP 2018 website: https://blackboxnlp.github.io/2018/
    BlackboxNLP 2018 proceedings: https://aclanthology.info/events/ws-2018#W18-54
    BlackboxNLP 2019 website: https://blackboxnlp.github.io/
    32 min
  • 80 - Leaderboards and Science, with Siva Reddy
    Originally used to entice fierce competitions in arcade games, leaderboards recently made their way into NLP research circles. Leaderboards could help mitigate some of the problems in how researchers run experiments and share results (e.g., accidentally overfitting models on a test set), but they also introduce new problems (e.g., breaking author anonymity in peer reviewing). In this episode, Siva Reddy joins us to talk about the good, the bad, and the ugly of using leaderboards in science. We also discuss potential solutions to address some of the outstanding problems with existing leaderboard frameworks.
    Software platforms for leaderboards:
    http://codalab.org/
    https://leaderboard.allenai.org/
    30 min
  • 79 - The glass ceiling in NLP, with Natalie Schluter
    In this episode, Natalie Schluter talks to us about a data-driven analysis of career progression of male vs. female researchers in NLP through the lens of mentor-mentee networks based on ~20K papers in the ACL anthology. Directed edges in the network describe a mentorship relation from the last author on a paper to the last author, and author names were annotated for gender when possible. Interesting observations include the increase of percentage of mentors (regardless of gender), and an increasing gap between the fraction of mentors who are males and females since the early 2000s. By analyzing the number of years between a researcher’s first publication and the year at which they achieve mentorship status at threshold T, defined by publishing T or more papers as a last author, Natalie also found that female researchers tend to take much longer to be mentors. Another interesting finding is that in-gender mentorship is a strong predictor of the mentee’s success in becoming mentors themselves. Finally, Natalie describes the bias preferential attachment model of Avin et al. (2015) and applies it to the gender-annotated mentor-mentee network in NLP, formally describing a glass ceiling in NLP for female researchers.
    https://www.semanticscholar.org/paper/The-glass-ceiling-in-NLP-Schluter/abfb1eb2d27194269503afce8be45909c8f86f4b
    See also:
    Homophily and the glass ceiling effect in social networks, at ITCS 2015, by Chen Avin, Barbara Keller, Zvi Lotker, Claire Mathieu, David Peleg, and Yvonne-Anne Pignolet. https://www.semanticscholar.org/paper/Homophily-and-the-Glass-Ceiling-Effect-in-Social-Avin-Keller/23dcb12dd918fcf29f7abb287dd466478031b8ff
    Apologies for the relatively poor audio quality on this one; we did our best.
    27 min
  • 78. Where do corpora come from?, with Matt Honnibal and Ines Montani
    Most NLP projects rely crucially on the quality of annotations used for training and evaluating models. In this episode, Matt and Ines of Explosion AI tell us how Prodigy can improve data annotation and model development workflows. Prodigy is an annotation tool implemented as a python library, and it comes with a web application and a command line interface. A developer can define input data streams and design simple annotation interfaces. Prodigy can help break down complex annotation decisions into a series of binary decisions, and it provides easy integration with spaCy models. Developers can specify how models should be modified as new annotations come in in an active learning framework.
    Prodigy: https://prodi.gy
    Prodigy recipe scripts: https://github.com/explosion/prodigy-recipes
    Twitter:
    https://twitter.com/_inesmontani
    https://twitter.com/honnibal
    31 min
  • 77. On Writing Quality Peer Reviews, with Noah A. Smith
    It's not uncommon for authors to be frustrated with the quality of peer reviews they receive in (NLP) conferences. In this episode, Noah A. Smith shares his advice on how to write good peer reviews. The structure Noah recommends for writing a peer review starts with a dispassionate summary of what a paper has to offer, followed by the strongest reasons the paper may be accepted, followed by the strongest reasons it may be rejected, and concludes with a list of minor, easy-to-fix problems (e.g., typos) which can be easily addressed in the camera ready. Noah stresses on the importance of thinking about how the reviews we write could demoralize (junior) researchers, and how to be precise and detailed when discussing the weaknesses of a paper to help the authors see the path forward. Other questions we discuss in this episode include: How to read a paper for reviewing purposes? How long it takes to review a paper and how many papers to review? What types of mistakes to be on the lookout for while reviewing? How to review pre-published work?
    39 min
  • 76 - Increasing In-Class Similarity by Retrofitting Embeddings with Demographics, with Dirk Hovy
    EMNLP 2018 paper by Dirk Hovy and Tommaso Fornaciari. https://www.semanticscholar.org/paper/Improving-Author-Attribute-Prediction-by-Linguistic-Hovy-Fornaciari/71aad8919c864f73108aafd8e926d44e9df51615
    In this episode, Dirk Hovy talks about natural language as social phenomenon which can provide insights about those who generate it. For example, this paper uses retrofitted embeddings to improve on two tasks: predicting the gender and age group of a person based on their online reviews. In this approach, authors embeddings are first generated using Doc2Vec, then retrofitted such that authors with similar attributes are closer in the vector space. In order to estimate the retrofitted vectors for authors with unknown attributes, a linear transformation is learned which maps Doc2Vec vectors to the retrofitted vectors. Dirk also used a similar approach to encode geographic information to model regional linguistic variations, in another EMNLP 2018 paper with Christoph Purschke titled “Capturing Regional Variation with Distributed Place Representations and Geographic Retrofitting” [link: https://www.semanticscholar.org/paper/Capturing-Regional-Variation-with-Distributed-Place-Hovy-Purschke/6d9babd835d0cdaaf175f098bb4fd61fd75b1be0].
    30 min
  • 75 - Reinforcement / Imitation Learning in NLP, with Hal Daumé III
    In this episode, we invite Hal Daumé to continue the discussion on reinforcement learning, focusing on how it has been used in NLP. We discuss how to reduce NLP problems into the reinforcement learning framework, and circumstances where it may or may not be useful. We discuss imitation learning, roll-in and roll-out, and how to approximate an expert with a reference policy.
    DAgger: https://www.semanticscholar.org/paper/A-Reduction-of-Imitation-Learning-and-Structured-to-Ross-Gordon/17eddf33b513ae1134abadab728bdbf6abab2a05?navId=citing-papers
    RESLOPE: http://legacydirs.umiacs.umd.edu/~hal/docs/daume18reslope.pdf
    44 min

About NLP Highlights

From the publisher's feed

**The podcast is currently on hiatus. For more active NLP content, check out the Holistic Intelligence Podcast linked below.**