get in touch Search
open menu close menu

How we built a grounded AI assistant for public libraries

The needs of public libraries

Public libraries have a strange kind of data challenge. The libraries have rich digital catalogs, and their users ask natural questions. Yet most discovery tools still respond as if it’s 2010, meaning a search box that provides a wall of carousel results and offers no room for a follow-up question.

We spent four months with a discovery and engagement platform serving public libraries across the United States and Canada, testing whether that could change. Here is how we built the proof of concept, what we learned, and where it is headed next.

 

The challenge of static discovery

Our client’s platform helps library patrons find books, events, and branch hours across hundreds of public libraries. But the experience stopped at search: type a query, scroll a carousel, and done. Anything more conversational, like ‘is this available near me,’ or ‘what’s happening this weekend for children,’ was completely missing from the platform.

On the staff side, there was also an informational gap. Every month, libraries collected hundreds of visitor reviews with real signals buried inside them, and no efficient way to read all of them, let alone act on them quickly.

And underneath it all was the difficult challenge of trust: could a large language model answer questions about a public institution’s data accurately, without inventing details or exposing anything sensitive?

 

The approach: grounding AI in real library data

We designed the assistant around one non-negotiable rule: it answers only with approved, library-controlled data: no general web knowledge, no guessing, no personal patron data in the loop at any point.

The architecture paired AWS Bedrock, running Anthropic’s Claude models, with a Bedrock Knowledge Base for retrieval. This grounded every answer in library-approved content rather than the open web. Using a managed model service instead of standing up our own meant the team could get straight to the harder challenge: making the answers actually reliable. A React chat interface sat atop a Python/FastAPI backend that orchestrated the underlying AI agents, with structured UI cards to display retrieved information in a friendlier manner, such as details of a found book, event, or branch location.

 

Multiple specialized agents, one grounded answer

Underneath the chat interface, the assistant isn’t one model answering everything. An intent-classification agent reads each question first and routes it to a specialized agent: one for book search and availability, one for events and branch locations, one for general website and FAQ content, and one that summarizes patron reviews into themes and sentiment. A conversation manager keeps track of context across turns, and anything genuinely ambiguous falls through to a general-purpose chat agent rather than getting a confident, yet incorrect answer.

Every specialized agent is restricted to its own approved data source, meaning that no agent answers from open-ended model knowledge, only from what the library has actually published.

How the pieces fit together

Behind those agents sits a straightforward data layer. The library-approved FAQs and resource content that the assistant may draw from are stored in Amazon S3.  In addition, Amazon Aurora covers the relational side, events, and locations, with no one on the team managing a database server. Amazon Redshift handles fast, structured querying over book-mapping data at a scale a simple lookup table wouldn’t handle well. And Amazon DynamoDB caches recent responses, which is a large part of why the assistant feels quick rather than “AI-slow.”

None of these are exotic choices. Each one is a mature, managed service doing the one job it has been built for, wired together rather than custom-built from scratch.

The result: faster, grounded, and ready to hand off

Performance work in the final phase brought the assistant’s average response time down significantly:

We treated security and governance as part of the deliverable, not an afterthought. IP allowlisting and controlled-environment access were in place before handover, and the system was built from day one never to process personal patron data.

By the end of the engagement, the assistant was documented, security-hardened, and handed over in full to our client’s engineering team, along with the sprint-by-sprint rationale for every major decision.

 

What this opens up next

The proof of concept did what it was designed to do: prove that generative AI could work inside a public-sector, multi-tenant environment without compromising on grounding, security, or governance. Because much of the stack – model access, retrieval, storage, caching — runs on managed cloud services rather than something we operate ourselves, extending it to more libraries looks more like configuration than a rebuild.

That validation, not just the working prototype, is what makes the architecture worth extending across a much larger base of public libraries.

Want to build something similar?

Andrei Scutariu,
Delivery Manager @ Yonder

Banner by Unsplash Marcel Strauss

STAY TUNED

Subscribe to our newsletter today and get regular updates on customer cases, blog posts, best practices and events.

subscribe