get in touch Search
open menu close menu

How AI reduced manual data entry and operational costs

In 2024, we were already integrating AI into solutions that automated data entry and reduced operational costs for one of our partners, which was focused on expansion. The articles shared recently on our insights page are just a continuation of that journey. 

But behind all these new solutions, the starting point is not AI itself but a business problem that becomes impossible to ignore as companies grow. Now, let us dive deep into how we built that architecture and the autonomous decision flow. 

 

The challenge of scaling manual work

Expanding into new markets is challenging, especially in regulated industries where each country introduces its own rules and documentation requirements. 

For our client, an international SaaS platform serving EHS (Environmental, Health, and Safety) and ESG (Environmental, Social, and Governance) compliance teams, much of the required information comes from third parties. Teams have to search for documents online, review them, and manually enter the data into internal systems. 

As the company expanded into new territories, that process was difficult to scale. Every new market meant repeating the same slow, manual steps. 

 

The approach of designing an AI-powered extraction pipeline

Yonder and the client’s team worked closely to design and implement a pipeline to automate document extraction. Our input in this partnership covered the full scope of the solution: from the machine learning workflow and software architecture design to product management and the integration of the pipeline into existing applications. Every output, from architecture decisions to integration points, went through the client’s team before it was shipped.  

A cross-functional team developed the extraction component using Python and large language models (LLMs). The LLM infrastructure was hosted across both Azure and AWS environments. The AI pipeline was integrated with two PHP applications and one .NET system that supported the broader operational workflow. 

To ensure transparency and control over the process, we implemented monitoring and governance mechanisms using Elasticsearch and Kibana. These tools allow the team to track pipeline performance, monitor data quality, and ensure proper governance of the extracted information. 

Yonder diagram pipeline architecture AI document extraction

 

From manual processing to autonomous document extraction

The first phase focused on a semi-autonomous implementation. The AI pipeline handled document extraction while human operators validated the results before the information entered the system. 

Even in this initial stage, the impact was significant. Data entry productivity increased nearly threefold within months after release, reducing the cost per processed document and lowering overall operational expenses. As the system stabilized, the required operational team size was reduced as well. 

Six months after the initial release, the project entered the second phase. During this period, we continued analyzing production data and refining the extraction process to improve accuracy and reliability. 

By the end of the second phase, the pipeline ran on its ownThe system extracted data from each document and then scored the quality of the extraction. Both the result and the confidence scores went to a rules engine, which ran a second layer of validation in code and made one call: does a human need to look at this, or not? Once the data cleared that check, it flowed automatically into the rest of the application ecosystem. 

Yonder diagram autonomous decision flow

 

The result of running the autonomous pipeline on productivity

Within months of releasing the autonomous version of the pipeline, productivity increased by more than fivefold. 

Thanks to the system’s high accuracy, about half of the processed documents were subsequently automatically published without human intervention. This significantly reduced the cost per document and lowered the operational effort required to manage the process. 

It also opened a door the old approach couldn’t. Because the LLMs handled multiple languages natively, the pipeline could move into new markets without a rewrite for each. The same system that lowered costs at home is what made expansion cheaper. 

Yonder diagram results AI document extraction results

A data-driven approach to continuous improvement 

Every improvement in the pipeline was guided by data rather than intuition. When a quality issue appeared, we analyzed which parameters were affected and how frequently the problem occurred. This allowed us to estimate the potential impact of each change before implementing it and forecast the expected gains. 

The approach helped the team continuously refine the system based on measurable outcomes. 

We applied a similar pipeline last year for another partner in a different industry, and the results again confirmed the value of automating document extraction. 

Years of experimenting with automation and AI have shaped how we approach these problems today, allowing us to design solutions that improve operational efficiency in practical, measurable ways. 

Yonder Diagram data driven loop

 

One of the most important aspects is starting with what we want to solve. Vlad, our Software Architect on this project, emphasized the best:

You need a real problem first. Then a team that knows the domain, wants to keep evolving, and won’t ship anything without a way to measure success. And they need the discipline to follow it through to the end, while staying conscious of time, cost, and the waves the industry is riding. If you have all that, an LLM is just another tool for digitalizing the world. 

Manual document work doesn’t scale. We’ve proven that fixing it does, across multiple industries. Let’s talk about your case.

By Vlad Cantor
Software Architect @ Yonder

STAY TUNED

Subscribe to our newsletter today and get regular updates on customer cases, blog posts, best practices and events.

subscribe