Private AI in Production: Four Deployments, One Architecture
The architecture page shows four processing paths — RAG, OCR, Voice Assistant, AI Crawler. This page is the proof: one real, in-production deployment on each path, pulled from our own published case studies, not written for this page.
Every deployment below runs on privately hosted models. No client's documents, tickets, test results, call audio or market data has been sent to a public model API — the same architecture and the same commitment described on the rest of this cluster, already shipped.
A support desk that answers in its own product's language
A CRM company in Europe needed real answers to specific product questions — without any ticket content leaving an environment they control.
Their support team handles questions about how the product is configured and behaves — permissions, integrations, data model, edge cases in workflow automation. Volume was routine and repetitive, but the questions were not simple: a useful answer often depended on how a specific account was set up.
Support tickets contain their customers' customer records. Under their own contractual commitments, that content could not be sent to a third-party model provider, whatever that provider's retention policy said — which ruled out most of the market before the project could start.
Product documentation and resolved-ticket history, so answers reflect how the product actually behaves, with the source retrievable.
The agent can look up account state and raise or update a ticket, so it completes the task rather than describing what to do next.
When the agent hands over, the conversation and the attempted resolution go with it.
Defined limits on what it will attempt, and an instruction to say it does not know rather than guess.
Entirely on privately hosted open-weight models. No conversation, ticket or account record is sent to OpenAI, Anthropic, Google or any other model provider — every request and response is logged inside the approved environment.
The privacy position is what made the project approvable at all
A build on a public model API would not have passed the client's own commitments to their customers. The choice was not between a cheaper approach and a more private one — it was between a private deployment and no project.
Test results that stopped being retyped
Results arrived as images and were keyed in by hand. We automated the read — and kept every figure traceable to the page it came from, because accreditation requires it.
Samples arrive from producers and retailers, are analysed, and results are reported back. The lab works to accreditation requirements — every figure it reports has to be traceable to its origin. Results came off instruments as images, and someone read those images and typed the values into the lab system by hand.
The obvious fix — automated extraction — was also the obvious risk: a pipeline that is usually right creates a review burden and an accreditation problem at the same time.
Values and tables read from images whose layout varies by instrument, with position on the page retained alongside each field.
Range checks and consistency rules catch what a model cannot know — a value a test would never produce.
Every extracted value carries a score. Low-confidence values never enter the lab system silently.
The source image is retained against every figure, with a record of what was extracted, flagged and confirmed — this is what made the automation acceptable.
On privately hosted models. No test data, sample reference or client identifier leaves the approved environment — deciding this first shaped the architecture rather than constraining it late.
In regulated work, provenance is not a feature. It is the requirement.
An accredited lab has to show where every number came from. The retained source image, field-level confidence and review record are what turned automation into something the lab could put its name to.
A voice agent that conducts screening interviews
We needed to screen candidates at volume without losing signal, so we built it for ourselves. It runs in production on our own hiring.
Hiring engineers means a lot of first-round calls — structured, necessary, and expensive in interviewer time. Volume meant calls competed with client work, so candidates waited. And because different people ran them, the notes varied: comparing two candidates meant comparing two different conversations.
Screening at scale without losing signal is exactly the shape of problem we tell clients to bring us. It seemed reasonable to solve it on ourselves first.
Not a form read aloud — the agent follows up on a vague answer and copes with someone talking over it.
Scored against criteria set for the role, with the transcript attached so any judgment can be checked.
The call happens when it suits them, which removed the scheduling delay entirely.
Vaani rejects nobody. It produces the first pass; the hiring decision stays with people — which the EU AI Act requires of recruitment systems.
Disclosure to the candidate at the start of the call, no automated rejection, transcript and scoring retained together, processing on private models with residency and retention set by whoever deploys it — required by the EU AI Act's high-risk classification for recruitment.
Voice fails in production, not in the demo
Every voice agent sounds good in a scripted demonstration. The difference shows on a bad line, with an interruption, at the fiftieth call of the day. This is the proof we put a voice agent into production, not a slide about one.
Market intelligence across multiple markets, on one method
A global financial institution needed a consistent read on several markets at once. Coverage was uneven, assembled by different teams, and hard to compare across regions.
The client operates across many countries and needed a current view of conditions in several markets at once. That work was already happening, but locally — different teams gathered different things, at different intervals, to different standards. Each market's picture was defensible on its own and almost impossible to place next to another.
The underlying problem was not access to data. It was the absence of one method applied consistently, with a record of where each figure came from.
The same definitions, sources and collection cadence applied across regions, so the picture can be read as a whole.
Each value carries where it came from and when it was collected — an analyst who cannot cite the source cannot use the number.
Values are graded rather than presented as uniformly reliable, so attention goes where the data is weakest.
Sources change and pages move — we run the pipeline and watch it, which is what makes the data current enough to act on.
Collection is scoped to what the client is permitted to gather and use, documented for their own compliance function. Processing runs on private models within the approved environment — nothing is routed to a third-party model provider.
Comparability was the deliverable
The client had plenty of data already. What they lacked was one method applied everywhere, so a difference between two markets could be read as real rather than an artefact of how each team collected it.
From Proof Back to the Architecture
These four deployments are what the rest of this cluster describes in the abstract — read backward from here, or start again from the architecture.
Your deployment could be the fifth one on this page.
Bring us the workflow, the document, the call or the market read your team handles by hand. We'll tell you honestly whether it's a private-AI problem — and which of these four paths it runs on.