
Sign up to save your podcasts
Or


Join us as we explore the power of ChatGPT and its impact on document analysis in the Apogee Suite. Our expert guests will delve into the different models of ChatGPT and how they can enhance creativity and language understanding. Discover how ChatGPT can help you ask the right questions for your document analysis process and why context is so critical in this field. Tune in to our podcast for a deep dive into the world of ChatGPT and document analysis.
Welcome to our latest episode of the podcast where we'll be delving into the world of NLU in Document AI.
We'll start by discussing Optical Character Recognition (OCR) and how it differs from Natural Language Understanding (NLU). OCR allows for the transformation of non-digital formats into digital ones, but it does not provide understanding of the content. NLU, on the other hand, goes beyond OCR by providing a readable version for the computer, allowing it to understand any kind of document.
Traditionally, understanding documents was a rule-based process that relied on the word order in a phrase. However, this system lacked semantic understanding and was not portable, requiring the building of new systems for each industry, resulting in longer and less accurate processes.
At 1000ml, we've worked with a variety of industries to adapt Contract AI, a field of Document AI, to their specific needs. For example, in the pharmaceutical industry, we built a pipeline for categorizing invoices and contracts according to the medicine they bought. In the legal industry, we worked on a system that allowed NLU to focus on patent documentation. In the financial and insurance industries, our work centered around form information extraction, resulting in faster and more accurate processes. And in healthcare, we combined different information about a patient in one place.
All paper-driven industries can benefit from Document AI, especially when NLU is involved. Join us next week as we continue to explore the exciting world of NLU in Document AI.
We started the year talking about NLP, and to continue it’s mandatory to talk about context, the contextual analysis of a text. Contextual analysis is understanding where a sentence is coming from.
When we think about how computers understand human language, we have to understand that it's a process and not an easy one also each language has a sense and syntax model that defines each one. So in order to process and understand that information, computers need signals to guide the process and to associate the language to action themselves.
We're at the tail end of the year now, and we're just talking about what 2023 may hold for 1000ml and generally in the world of NLP.
So 1000ml has a suite of NLP and AI products that we called Apogee. The suite is comprised of just general text ingestion that can be documents, which is usually what people go for but also allows you to ingest any website or webpage basically whatever you want that has text in it, including.
The crux of it is that Apogee suite builds up a series of pipelines and APIs including a recommendation engine and relevancy engine API that allows all the tools that we build on top of it to utilize these engines to properly search for content. So you might think of things like, okay, well, I go online and I do a Google search and I'm looking for a flight or something.
Imagine all kinds of descriptions, text messages, and emails that you send internally, giving you the ability to ask content questions. So imagine a use case like you're in e-commerce and you're looking for the right type of pants. Somebody's like, I need black pants, you can go all the way from like athleisure to like super serious, like tuxedo pants if you're a man. So are you just gonna surface everything along the lines and how do you know exactly the kind of things that they're looking for?
ADA sort of allows you to ask your database, your text, anything. The curated data searching that we've done prior with just our semantic search is now getting a big boost, allowing us the ability to really ask the document things, so not just metadata things like who's the author, which basically any system could do.
When will 1000ml tackle as ambitious Witness Prep AI, come out of the world of us doing a lot of work in legal and judicial and deeply understanding the paradigms that exist there with the advent of us creating semantic search and then layering on ADAs so you can ask the content anything and you could read any content. So again, Witness Prep AI is all about ingesting a lot of things about a specific case or a specific area of law or kind of context that may occur.
It's largely going to be an exercise for the enterprise world until we really understand how to maximize its value and use and then can reduce it to a point where it's then manageable for people who want it as an add-on into our main suite, whether that's for our regular B2C clients or for clients and customers.
So that's a bit of a view into 2023 and what 1000ml is gonna be working on. It's still all about NLP and AI and mostly focused on Apogee Suite. It's our big framework and it's what we sell the most of.
2022 has been a fun year in AI and NLP, and today we thought we'd take a minute and reflect back on all the things that have occurred. Some of the important things that have occurred in the world of NLP and a lot of the generative world of NLP in 2022.
GPT has been a big deal in NLP, Generative, Pre-trained Transformer is a generation technology to help us generate usually text content and that can really be like all kinds of things. If you go look on YouTube and you kind of like dabble around with GPT-3 OpenAI, you'll see examples of all kinds all the way from like, give me a tweet basically, you know, like generate a tweet from you, which is kind of interest.
So for a machine to generate and do it well and in within the context of what you want is actually really cool and really good. So this year, obviously there's been advances in GPT-3 they did allow for the editing and insertion of data, like human data back into the reports. So if you were to ask GPT to specifically generate something for you, could actually embed new information in there to make it slightly better is a big deal.
Another thing that they've done since they noticed in the past that was happening is they've added a lot of plagiarism detection into their generation. They've done that with slightly better word parsing and quite a bit of synonym generation so that the content will be as unique as possible and avoid plagiarism.
Another really big thing done this year was to allow for the monitoring or ongoing review of data repositories that allow GPT to constantly scan and add new content as it becomes available and then make GPT's engine more relevant to your search, as well.
Tune in next week we're gonna talk a little bit about our roadmap for 2023 and what we see could happen in NLP next year.
Happy 2023 everyone!
We started to wrap up the year a little bit talking about the considerations that you have to think of when you're deciding whether to build or buy AI. That really feeds in well and quite nicely into today's topic where we're gonna talk more about the technical knowledge required for AI projects. So this is assuming you went through the framework of deciding whether to build or buy, and you were like, you know what? We can build this, let's do this ourselves.
So now that you're in the world of let's build this because you've decided that we have the staffing resources, and we're able to get consultants who could do this, having enough know-how and enough people with the breadth of knowledge required for this. Also, we've worked with or have the infrastructure available for this.
Generally, the organization is in a good shape, or our change management processes are quite good, understanding how to make sure that this blends into our organization well having in mind the total cost of ownership of such a project.
For a very long time anybody who's done a lot of work in NLP has chosen to usually start their journey with the package NLTK is largely like the go-to and it is the basis, it's the foundation that provides quite a bit of functionality, but some of the things that are critical for you to know if you're going to do serious NLP you need to know stemming.
You need to split the content either into phrases because you may want to analyze phrases or sentences, or split it into paragraphs, but usually into words in order to understand and know how to use a part of speech triggers bringing you endless opportunities.
For example, there's the possibility of doing a strictly extractive summary where you're going to pull things out of a document. Imagine an entire thesis, there is an abstract, and using an NLP on that whole thesis, you'd probably just pull that abstract becoming your summary.
.So you want to make an abstract, so you're largely looking to take the knowledge and understand it and then abstract the knowledge of that document. So you're really summarizing as opposed to pulling specific parts of the content out.
Again, high risk for it cause we're really good at this stuff. But you could also use machine learning and AI as a sort of input to NLP. You'd find that there are many clusters when you're doing this that can then help lead you to create an AI program based on those clusters.
Lastly, if you wanted to use the NLP as an input to different programs, you could do that for things like the things we do, which we do a lot of work in document intelligence, in contract AI.
Today we're gonna talk about whether to buy or build AI systems.
As opposed to the shorter write-ups that we usually have, you'll probably get a lot of benefit from rating it and even using it as notes when you are making the decision on whether to build or buy a new system generally, not just AI.
Like building or buying most technology systems or new systems for our company you have to start thinking of the total cost of ownership. That's a big deal for people for companies actually and a lot is wrapped up in that. At the end of the day though the biggest hurdle that most companies have when deciding whether or not to build or buy a system is that there needs to be a stakeholder who is going to really be responsible for it.
So generally in a buying scenario, and you see this a lot for big organizations and government entities the smart play is to understand all the features that your company needs. It could be a big system or a little system, you should come up with a matrix of sorts or a checklist. You want really good and clean user interfaces and user experience, an easy example is Google.
You also want to think of the ability to change the data model or AI model, and whether that's baked into the app and you know, rigid or whether you can actually affect those models that could and should be important to you.
Think that technology does 85% of what you want, and you're okay with that, but you can create a roadmap to get to closer to a hundred percent down the future, not costing anything. Finally, you want to think of ecosystems, like communities, the number of partners they have all the users, are people writing about that. Are there examples and forms and stack overflow in all these places do they have a large list of clients and do they provide training? Is there self-led training that they provide? Do they come in the house and do some training? So there's all of that for thinking of future proof in your purchase.
We've been talking a lot about our witness prep AI and what's gone into making all of that and today we thought we'd actually culminate into its intended uses and a bit about how we deliver it.
It's possible to acquire all the data that you need especially about cases and laws and decisions and whatnot you have to build a language model. There's been an academic language model called FILAC and we've extended that.
So given the ability to now understand legal texts and to model how to get to an outcome, basically to have all of them, I can render a decision, a law an actual case, whatever into a computer, a usable bit of data so that the computer can actually ingest it properly and not just have the unstructured text.
So now that you have this AI model, how do you deliver that to the best? You have to think about what the most natural thing for people is, when you're doing a user interface exercise you don't want to take people to a place where they have to think about the UI while they're also thinking about the context and problem that you're working on.
Because then your brain kind of fractures itself and you're lost of it. So a very common use case for most people just searches for a really good job of search and curation, which is something that we do.
It's all in the end powered by the AI model. But the way to deliver that is a very curated search especially if you're able to do it in a hierarchical view of what you're searching for so that you have, like a meta topic and you get yourself all the way down to specifics that get people the idea.
Most law offices & legal professionals go through common or similar cases, when they work on a case they look for prior arguments, prior cases, and precedents, allowing an automatization process to be implemented since it makes it searchable and usable information.
They use a sort of search tool and it's generally only a keyword search tool, since someone with a lot of experience, will know that in each case what to look at, the specific set of keywords that is pretty unique and doesn't usually happen in the rest of them. Of course, that other hand has some trouble since limiting to only those keywords you could lose relevant other information.
So when you need to know the exact keywords in order to search for precedent and prior cases. Imagine that there is the need to understand all of these details and information to really build a case up so that it's possible to understand what people have said in the past, what's worked and what hasn't, these being the material facts of the case.
The other thing that you really can't do without actual live debates and trials is to judge the strength of the argument, so this type of judging will obviously going to rely on the entirety of that package. In order to really understand and get to a point where there is confidence in an argument it really benefits everybody to, have sort of a score about it, so when there is an argument with a high degree of confidence, that there is the conviction that will or power up that case over the other side.
There's an opportunity here to make it more scientific and that's where we shine quite a lot in getting into the AI for legal decisions of courts and obviously also pushing the envelope and innovation in the world of witness readiness and witness prep.
In order to get to a world where you can build an AI program you're going to have to build up your internal capabilities and knowledge because you're dealing largely with unstructured texts, for example, images and video, those are definitely unstructured, but you can structure them by pre-processing in text. You can do the same, you can turn extreme, extremely large pieces of content, whether they are legal decisions.
From the publisher's feed