Assistants that pick up where they left off
A good human assistant never asks how you take your coffee twice. Most AI assistants ask every session, because what you told them lived in a context window that has since closed.
The mixture
- 01 · Persistent memory
- Leads
- 02 · Current context
- Light
- 03 · Traceable decisions
- Uses
- 04 · Shared context
- Uses
- 05 · Scoped retrieval
- Uses
Every session starts from zero
Not "the model forgets." Three specific mechanisms, each of which you can find in your own transcripts this afternoon.
- 01
Corrections do not stick
The user said "never book before 10am" last Tuesday. The agent complied, and the session ended. Nothing wrote the rule down, so this Tuesday it books a 9am and the user corrects it again.
- 02
The model swap resets the relationship
You moved to a cheaper model. The preferences lived inside the old provider's built-in memory, so the switch quietly wiped months of accumulated context along with it.
- 03
History is stuffed, not selected
To compensate, the whole chat log goes into the prompt. Cost climbs, latency climbs, and the one line that mattered is buried under three hundred that did not.
Mostly persistent memory. Then three others.
Every application on Alchemyst is a different mixture of the same five jobs. That mixture is what makes this a different piece of software from the one next to it, even though the API underneath is identical.
What the user told you once stays told
Memory is written per user and per session with context.memory.add, and updated rather than duplicated. It lives in your context layer, not inside any one model, so it follows the user across sessions, devices and providers.
- Preferences return on the first turn, not the fifth.
- Corrections overwrite, they do not pile up.
- Swapping models does not reset the user.
Uses · Shared context
Every surface, one user
The email assistant and the calendar assistant read the same memory, scoped to the same person.
Uses · Scoped retrieval
Only the relevant preference
Travel preferences stay out of the expense report. Scopes decide what the window sees.
Uses · Traceable decisions
Why it did that
When the assistant acts on a remembered preference, the trace shows which memory it used, and from when.
That is four of the five. The fifth, current context (superseded versions are subtracted before the model ever sees them), is what leads in Employee support, Contract & legal ops and Financial ops instead. Same API, different mixture.
The sources you already have
The assistant already sees the conversation. Write the durable parts back as memory, and pull the rest from where it already lives. The Vercel AI SDK middleware and LangChain integrations do the wiring for you.
01 · Ingest
Scope what you ingest
Every document lands with a groupName: the sets it belongs to. Those sets are what retrieval intersects later, so the structure you choose here is the precision you get there.
context.addimport AlchemystAI from "@alchemystai/sdk";const client = new AlchemystAI(); // reads ALCHEMYST_AI_API_KEYawait client.v1.context.add({ context_type: "resource", scope: "internal", source: "google-docs", documents: [{ content: "Travel: aisle seat, no red-eye flights, hotels within 1km of the meeting venue.", }], metadata: { fileName: "travel-preferences.md", groupName: ["assistant", "user_4821"], // the sets this belongs to },});02 · Write
Write what the user corrected
The transcript is not the memory. The correction is. Write it once, against the user, and it is there next time.
context.memory.addawait client.v1.context.memory.add({ sessionId: "user_4821", contents: [ { role: "user", content: "Please never book anything before 10am." }, { role: "assistant", content: "Understood. All bookings at 10am or later." }, ], metadata: { groupName: ["assistant", "user_4821"] },});03 · Search
Search before replying
Search intersects the scopes, subtracts superseded and duplicate content, and ranks what survives. Only that reaches the model, and the whole decision is recorded as a Context Trace.
context.searchconst { contexts } = await client.v1.context.search({ query: "Book a meeting with the Acme team next week", scope: "internal", similarity_threshold: 0.8, minimum_similarity_threshold: 0.5, metadata: { groupName: ["assistant", "user_4821"] }, // ∩ narrow scope});// − superseded, deduplicated → ranked → into the window// Every search is recorded as a Context Trace.const reply = await llm.respond(message, { context: contexts });
npm install @alchemystai/sdk or pip install alchemystai, both ship the same client. Full reference in the docs.
A user's memory is their data, not your vendor's
Preferences, habits and corrections are personal data under GDPR, the DPDP Act and CCPA. They need to be exportable and deletable on request, and they should not be locked inside a model provider's memory feature you cannot inspect.
Security & complianceManaged cloud
Encrypted in transit and at rest, isolated per organization, and scoped at write time.
Dedicated infrastructure
EnterpriseSingle-tenant, with VPC peering when your data cannot share a network boundary.
Self-hosted
On-premiseRun the context layer on your own infrastructure, with OpenTelemetry for observability.
The same API, a different mixture
Each of these leads with a different job, pulls from a different set of sources and needs a different call. All of them, by job.
Bring the assistant that keeps asking.
Not the demo. The one whose users have stopped correcting it because they have given up.