Voice agents get one round trip, and it has to be the right one
In a chat, a slow answer is a spinner. On a call, it is dead air, and the caller starts talking over the agent. Voice leaves no budget for retrieving the wrong thing.
The mixture
- 01 · Persistent memory
- Uses
- 02 · Current context
- Uses
- 03 · Traceable decisions
- Uses
- 04 · Shared context
- Light
- 05 · Scoped retrieval
- Leads
Latency is the whole product
Every retrieval choice that is invisible in text is audible on a call.
- 01
Every extra lookup is audible
Agents that search, reflect and search again feel fine in text. On a call, the second lookup is the moment the caller says "hello?"
- 02
Wide retrieval is slow retrieval
Searching the whole knowledge base for every utterance spends latency the conversation does not have, and most of what comes back is irrelevant to this caller.
- 03
Callers repeat themselves
The caller explained the problem yesterday. Today's agent asks again, and a caller who has to repeat themselves asks for a human.
Why voice AI teams choose Alchemyst
Verified capabilities built for enterprises deploying voice AI at scale.
Unmatched context speed
Complex conversation context is processed in 170ms, so voice agents respond instantly, without latency delays.
Superior context awareness
The highest memory F1 score in the industry (0.76) keeps conversation continuity precise.
Enterprise economics
Save 83% on costs while delivering 12x more performance value per dollar than competing solutions.
Verified and transparent
Tested in December 2025 on publicly available benchmarks. No hidden claims, only results.
Pareto frontier performance
Positioned at the efficiency frontier: the best balance of cost and performance available today.
Easy integration
REST APIs and SDKs that integrate with your existing voice stack in hours, not months.
170ms latency. Best-in-class efficiency.
Benchmark-proven performance for voice, tested in December 2025 on publicly available benchmarks. This is how the Alchemyst context engine defines the new Pareto frontier for voice AI.
- 170ms
- P50 latency
- 12x
- Value ratio
- 83%
- Cost savings
- 0.76
- Memory F1 score
Real-time voice AI responses
More intelligence per dollar
vs traditional engines
Superior context retention
| Metric | Competitors | Alchemyst |
|---|---|---|
| P50 latency | 500-800ms | 170ms |
| Memory F1 score | 0.45-0.58 | 0.76 |
| Value ratio (performance / cost) | 1x-2x | 12x |
| Cost savings | Baseline | 83% reduction |
| Setup time | 2-4 weeks | Under 24 hours |
Mostly scoped retrieval. Then three others.
Every application on Alchemyst is a different mixture of the same five jobs. That mixture is what makes this a different piece of software from the one next to it, even though the API underneath is identical.
The nearest match stops winning
Scopes are set before the call connects: this caller, this account, this intent. Each turn runs one narrow search in fast mode, so what comes back is small, relevant and inside the latency budget.
- One narrow search per turn.
- Scopes set before the first word.
- Fast mode for in-call lookups.
Uses · Persistent memory
Yesterday's call, remembered
What the caller said last time is written back and read before they speak.
Uses · Current context
Today's hours and offers
Superseded scripts and promotions are subtracted, so the agent never offers what has expired.
Uses · Traceable decisions
Why the agent said that
Every spoken answer carries the sources behind it, for QA review.
That is four of the five. The fifth, shared context (what one agent learns, the next one already knows, on the same definitions), is what leads in Sales agents, Customer success and Enterprise operations instead. Same API, different mixture.
Transform any voice use case
Deliver the same context engine across your entire voice AI stack.
Sales & outbound calling
Qualify leads and close deals with context-aware agents that understand customer history and objections.
3x faster call completionCustomer support
Resolve issues faster with instant context retrieval on customer history and preferences.
40% faster resolutionCollections & reminders
Intelligent due reminders that understand payment history and customer circumstances.
25% improvement in recovery ratesInbound call handling
Route and handle incoming calls intelligently, with complete context on the first ring.
Instant smart routing
The sources you already have
Works with your voice stack: Plivo, Twilio and any existing telephony, with OpenAI or Anthropic models. The LiveKit plugin adds persistent cross-session memory, and production-ready APIs deploy in under 24 hours.
01 · Ingest
Scope what you ingest
Every document lands with a groupName: the sets it belongs to. Those sets are what retrieval intersects later, so the structure you choose here is the precision you get there.
context.addimport AlchemystAI from "@alchemystai/sdk";const client = new AlchemystAI(); // reads ALCHEMYST_AI_API_KEYawait client.v1.context.add({ context_type: "resource", scope: "internal", source: "kb", documents: [{ content: "Standard delivery: 2 to 4 business days. Same-day delivery in Bengaluru and Mumbai for orders placed before 1pm.", }], metadata: { fileName: "delivery-windows.md", groupName: ["voice", "support", "delivery"], // the sets this belongs to },});02 · Write
Write what the caller needs next time
The call ends and the transcript is long. Keep the part the next call needs: the issue, what was promised, what is still open.
context.memory.addawait client.v1.context.memory.add({ sessionId: "caller_4410", contents: [{ role: "assistant", content: "Order #88213 reported missing. Promised a callback by 6pm with a courier update.", }], metadata: { groupName: ["voice", "support", "caller_4410"] },});03 · Search
Search before speaking
Search intersects the scopes, subtracts superseded and duplicate content, and ranks what survives. Only that reaches the model, and the whole decision is recorded as a Context Trace.
context.searchconst { contexts } = await client.v1.context.search({ query: "Where is my order?", scope: "internal", mode: "fast", similarity_threshold: 0.8, minimum_similarity_threshold: 0.5, metadata: { groupName: ["voice", "support", "delivery"] }, // ∩ narrow scope});// − superseded, deduplicated → ranked → into the window// Every search is recorded as a Context Trace.const reply = await llm.respond(message, { context: contexts });
npm install @alchemystai/sdk or pip install alchemystai, both ship the same client. Full reference in the docs.
Every network hop is audible
Voice agents feel the distance between the context layer and the model. Keep retrieval close to inference, keep scopes tight, and keep caller memory deletable when a customer asks.
Security & complianceManaged cloud
Encrypted in transit and at rest, isolated per organization, and scoped at write time.
Dedicated infrastructure
EnterpriseSingle-tenant, with VPC peering when your data cannot share a network boundary.
Self-hosted
On-premiseRun the context layer on your own infrastructure, with OpenTelemetry for observability.
Frequently asked questions
What makes Alchemyst different?
Alchemyst is the only AI context engine with verified, publicly available benchmarks. Our 170ms P50 latency and 0.76 memory F1 score are tested and proven, not marketing claims.
How much can we save?
Alchemyst delivers 12x more value per dollar with 83% cost savings versus traditional engines. A typical enterprise saves $50K to $500K a year, depending on call volume.
How quickly can we deploy?
In less than 24 hours. The APIs are designed for rapid integration with Twilio, Plivo or any custom telephony system.
Is there a setup cost or long-term contract?
No setup fees and no contracts. You only pay for what you use. Start free with a 30-day trial and full production access.
The same API, a different mixture
Each of these leads with a different job, pulls from a different set of sources and needs a different call. All of them, by job.
Bring the call with the dead air.
The one where the caller hung up before the agent found the answer.
No credit card required · Deploy in under 24 hours · 30-day free trial