NLP Highlights

NLP Highlights

By Allen Institute for Artificial IntelligenceScience
Download on the App Store

NLP Highlights episodes

  • 34 - Translating Neuralese, with Jacob Andreas
    ACL 2017 paper by Jacob Andreas, Anca D. Dragan, and Dan Klein.
    Jacob comes on to tell us about the paper. The paper focuses on multi-agent dialogue tasks, where two learning systems need to figure out a way to communicate with each other to solve some problem. These agents might be figuring out communication protocols that are very different from what humans would come up with in the same situation, and Jacob introduces some clever ways to figure out what the learned communication protocol looks like - you find human messages that induce the same beliefs in the listener as the robot messages. Jacob tells us about this work, and we conclude with a brief discussion of the more general issue of interpreting neural models.
    https://www.semanticscholar.org/paper/Translating-Neuralese-Andreas-Dragan/49612dc348ce953027bb4aba95adad0c703d76d1
    32 min
  • 33 - Entity Linking via Joint Encoding of Types, Descriptions, and Context, with Nitish Gupta
    EMNLP 2017 paper by Nitish Gupta, Sameer Singh, and Dan Roth.
    Nitish comes on to talk to us about his paper, which presents a new entity linking model that both unifies prior sources of information into a single neural model, and trains that model in a domain-agnostic way, so it can be transferred to new domains without much performance degradation.
    https://www.semanticscholar.org/paper/Entity-Linking-via-Joint-Encoding-of-Types-Descrip-Gupta-Singh/a66b6a3ac0aa9af6c178c1d1a4a97fd14a882353
    25 min
  • 32 - The Effect of Different Writing Tasks on Linguistic Style, with Roy Schwartz
    CoNLL 2017 paper, by Roy Schwartz, Maarten Sap, Ioannis Konstas, Leila Zilles, Yejin Choi, and Noah A. Smith.
    Roy comes on to talk to us about the paper. They analyzed the ROCStories corpus, which was created with three separate tasks on Mechanical Turk. They found that there were enough stylistic differences between the text generated from each task that they could get very good performance on the ROCStories cloze task just by looking at the style, ignoring the information you're supposed to use to solve the task. Roy talks to us about this finding, and about how hard it is to generate datasets that don't have some kind of flaw (hint: they all have problems).
    https://www.semanticscholar.org/paper/The-Effect-of-Different-Writing-Tasks-on-Linguisti-Schwartz-Sap/1a697d7cf187e51d5ccc23eb3ee5d2950ece5522
    25 min
  • 31 - Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
    ICLR 2017 paper by Hakan Inan, Khashayar Khosravi, Richard Socher, presented by Waleed.
    The paper presents some tricks for training better language models.
    It introduces a modified loss function for language modeling, where producing a word that is similar to the target word is not penalized as much as producing a word that is very different to the target (I've seen this in other places, e.g., image classification, but not in language modeling). They also give theoretical and empirical justification for tying input and output embeddings.
    https://www.semanticscholar.org/paper/Tying-Word-Vectors-and-Word-Classifiers-A-Loss-Fra-Inan-Khosravi/424aef7340ee618132cc3314669400e23ad910ba
    12 min
  • 30 - Probabilistic Typology: Deep Generative Models of Vowel Inventories
    Paper by Ryan Cotterell and Jason Eisner, presented by Matt.
    This paper won the best paper award at ACL 2017. It's also quite outside the typical focus areas that you see at NLP conferences, trying to build generative models of vowel vocabularies in languages. That means we give quite a bit of set up, to try to help someone not familiar with this area understand what's going on. That makes this episode quite a bit longer than a typical non-interview episode.
    https://www.semanticscholar.org/paper/Probabilistic-Typology-Deep-Generative-Models-of-V-Cotterell-Eisner/6fad97c4fe0cfb92478d8a17a4e6aaa8637d8222
    32 min
  • 29 - Neural machine translation via binary code prediction, with Graham Neubig
    ACL 2017 paper, by Yusuke Oda and others (including Graham Neubig) at Nara Institute of Science and Technology (Graham is now at Carnegie Mellon University).
    Graham comes on to talk to us about neural machine translation generally, and about this ACL paper in particular. We spend the first half of the episode talking about major milestones in neural machine translation and why it is so much more effective than previous methods (spoiler: stronger language models help a lot). We then talk about the specifics of binary code prediction, how it's related to a hierarchical or class-factored softmax, and how to make it robust to off-by-one-bit errors.
    Paper link: https://www.semanticscholar.org/paper/Neural-Machine-Translation-via-Binary-Code-Predict-Oda-Arthur/bbedfd0380eb2e62f1c3b61aaf484d5867e6358d
    An example of the Language log posts that we discussed: http://languagelog.ldc.upenn.edu/nll/?p=33613 (there are many more).
    39 min
  • 28 - Data Programming: Creating Large Training Sets, Quickly
    NIPS 2016 paper by Alexander Ratner and coauthors in Chris Ré's group at Stanford, presented by Waleed.
    The paper presents a method for generating labels for an unlabeled dataset by combining a number of weak labelers. This changes the annotation effort from looking at individual examples to constructing a large number of noisy labeling heuristics, a task the authors call "data programming". Then you learn a model that intelligently aggregates information from the weak labelers to create a weighted "supervised" training set. We talk about this method, how it works, how it's related to ideas like co-training, and when you might want to use it.
    https://www.semanticscholar.org/paper/Data-Programming-Creating-Large-Training-Sets-Quic-Ratner-Sa/37acbbbcfe9d8eb89e5b01da28dac6d44c3903ee
    26 min
  • 27 - What do Neural Machine Translation Models Learn about Morphology?, with Yonatan Belinkov
    ACL 2017 paper by Yonatan Belinkov and others at MIT and QCRI.
    Yonatan comes on to tell us about their work. They trained a neural MT system, then learned models on top of the NMT representation layers to do morphology tasks, trying to probe how much morphological information is encoded by the MT system. We talk about the specifics of their model and experiments, insights they got from doing these experiments, and how this work relates to other work on representation learning in NLP.
    https://www.semanticscholar.org/paper/What-do-Neural-Machine-Translation-Models-Learn-ab-Belinkov-Durrani/37ac87ccea1cc9c78a0921693dd3321246e5ef07
    30 min
  • 26 - Structured Attention Networks, with Yoon Kim
    ICLR 2017 paper, by Yoon Kim, Carl Denton, Luong Hoang, and Sasha Rush.
    Yoon comes on to talk with us about his paper. The paper shows how standard attentions can be seen as an expected feature count computation, and can be generalized to other kinds of expected feature counts, as long as we have efficient, differentiable algorithms for computing those marginals, like the forward-backward and inside-outside algorithms. We talk with Yoon about how this works, the experiments they ran to test this idea, and interesting implications of their work.
    https://www.semanticscholar.org/paper/Structured-Attention-Networks-Kim-Denton/0aec1745d0e054e8d86d21b20d0ee5fc0d932a49
    Yoon also brought up a more recent paper by Yang Liu and Mirella Lapata that computes a very similar kind of structured attention, but does so much more efficiently. That paper is here: https://www.semanticscholar.org/paper/Learning-Structured-Text-Representations-Liu-Lapata/4435c3586364e8f8a2c8c9ee671c39d7df7e196c.
    26 min
  • 25 - Neural Semantic Parsing over Multiple Knowledge-bases
    ACL 2017 short paper, by Jonathan Herzig and Jonathan Berant.
    This is a nice, obvious-in-hindsight paper that applies a frustratingly-easy-domain-adaptation-like approach to semantic parsing, similar to the multi-task semantic dependency parsing approach we talked to Noah Smith about recently. Because there is limited training data available for complex logical constructs (like argmax, or comparatives), but the mapping from language onto these constructions is typically constant across domains, domain adaptation can give a nice, though somewhat small, boost in performance.
    NB: I felt like I struggled a bit with describing this clearly. Not my best episode. Hopefully it's still useful.
    https://www.semanticscholar.org/paper/Neural-Semantic-Parsing-over-Multiple-Knowledge-ba-Herzig-Berant/6611cf821f589111adfc0a6fbb426fa726f4a9af
    11 min

About NLP Highlights

From the publisher's feed

**The podcast is currently on hiatus. For more active NLP content, check out the Holistic Intelligence Podcast linked below.**