Changing Embedding Models: Zengram Before and After Ptah
We tested a small Zengram integration that builds replacement embeddings with Ptah while old search keeps serving, then coordinates the final handover.
Ptah helps you plan, review, and apply database migrations. Try in your browser
A search request arrives while an application is changing its embedding model. The new model produces 384 numbers per document. The database still contains 1,536 numbers per document, calculated by the previous model. Changing an API setting cannot make those stored values belong to the new model.
For a small corpus, rebuilding everything during maintenance may be enough. For a corpus with millions of rows, that rebuild could become the longest part of the deployment. The search service needs a usable representation throughout that work, including when the replacement provider fails.
Zengram is a shared memory service for AI agents. It stores memories in PostgreSQL and retrieves them through vector, keyword, and entity-based search. Its embedding replacement script gave us a concrete place to test a different migration shape: build a new representation beside the active one, check it, and switch deliberately.
We implemented and tested that approach in an upstream proposal. The PR is open; this is a case study of our tested integration, not an announcement of Zengram adopting Ptah. The examples use Ptah 0.11.4, PostgreSQL 16.15, and pgvector 0.8.7. The supporting files contain the specification, test inputs, and recorded output.
The model is part of the stored data
Section titled “The model is part of the stored data”An embedding is a list of numbers that a model calculates from an input such as a sentence. Similar inputs should occupy nearby positions in that model’s vector space. A vector’s dimension is the length of the list: 1,536 values or 384 values in this example.
To search, the application embeds the query and compares that vector with stored document vectors. Both sides must use compatible coordinates. Imagine looking for “database connection failures” among stored troubleshooting notes: the query and each note need to be represented by the same model, with the appropriate query and document preprocessing.
Different dimensions make a mismatch obvious: PostgreSQL cannot compare the vectors as if their lengths matched. Equal dimensions do not make different models interchangeable. Two models can each return 384 values while assigning those coordinates different meanings. A comparison can then execute and still produce meaningless rankings.
PostgreSQL’s pgvector extension supplies vector storage and distance operators. An HNSW index builds a graph that helps find nearby vectors without comparing the query with every row. That graph belongs to the vectors it indexes. Changing the model therefore coordinates stored data, the search index, and the encoder configuration used by the application.
Before Ptah: rebuild the representation being searched
Section titled “Before Ptah: rebuild the representation being searched”Zengram’s api/scripts/reembed.js reads each memory’s persisted text and asks
the currently configured provider to embed it. Its destination is the same
memories.vector column that the API searches.
When the script detects a dimension change, it drops
idx_memories_vector_hnsw, changes the column to the new vector dimension using
NULL values, and then writes replacement embeddings row by row. At the end it
rebuilds HNSW. Payloads and the separate keyword-search data remain in place.
The script has useful operational controls: dry-run is the default, embedding calls have bounded concurrency, and a failed row is reported. It can be rerun. Those controls do not preserve the old vector representation once the dimension change has cleared it. A provider failure can leave some rows without vectors until a later attempt succeeds.
This is a manageable procedure when maintenance is short and its search impact is acceptable. As the corpus grows, completion time and provider availability become conditions for restoring the active representation. Zengram’s keyword fallback can still help answer requests, but it is not the old vector search.
A new column alone is not a finished migration
Section titled “A new column alone is not a finished migration”Building replacement vectors takes more than adding vector(384) to a table.
The provider can fail halfway through. A backfill—the pass that computes values
for existing rows—can take much longer than expected. A completed backfill can
still be missing a newly inserted memory, or contain the earlier text of a
memory edited while the provider was processing it.
Deletes matter too. An embedding request already in flight must not bring a deleted memory back into the candidate representation. The new index must be ready before it is relied upon. Document prefixes, truncation rules, and model selection must agree with the configuration that will produce future writes.
Finally, deploying the new query encoder before its matching stored vectors are ready creates a mismatch at the application boundary. This is a migration of data, model, index, and application configuration. The table alteration is only its starting point.
Build beside the active generation
Section titled “Build beside the active generation”A generation groups vectors produced with a particular model, preprocessing
configuration, and destination. In our integration, the original vectors stay
in memories.vector. The candidate goes into memories.embedding_v2.
The API keeps its existing encoder during this build. Ptah uses the replacement encoder to populate the candidate. A failed candidate request does not clear the original column or drop its index. It changes whether the replacement is ready, rather than removing the representation that already works.
Use a new column and run identifier for each model change, including changes that preserve dimensions. This keeps a retry of the same build distinct from starting a different generation.
Expose the text the application actually stores
Section titled “Expose the text the application actually stores”Zengram keeps source text inside a JSON payload. Ptah’s source specification
reads columns. The integration adds a stored generated embedding_text column
that follows the same text fallback as reembed.js: text, then content,
then note, then title.
-- Expose the same text fallback as reembed.js to Ptah's column-based source.-- A generated column follows payload edits without another application write.-- Adding a STORED column rewrites the table: schedule this one-time setup.SET lock_timeout = '5s';ALTER TABLE public.memories ADD COLUMN embedding_text text GENERATED ALWAYS AS ( COALESCE(NULLIF(payload->>'text', ''), NULLIF(payload->>'content', ''), NULLIF(payload->>'note', ''), payload->>'title', '')) STORED;The generated expression follows payload edits automatically. It avoids asking every application writer to maintain another copy of the text. This has a real setup cost: adding a stored generated column rewrites and locks the table. Schedule that one-time operation separately. The five-second lock timeout bounds waiting to acquire a lock, not the duration of the rewrite after acquisition.
Prepare and backfill the candidate
Section titled “Prepare and backfill the candidate”The complete specification selects public.memories,
uses id as its row key, reads embedding_text, and targets embedding_v2.
It declares a mutable source with outbox consistency, a cosine-distance HNSW
index, and a policy requiring explicit cutover approval.
The replacement endpoint in this worked example is a deterministic local test
provider. It returns 384-dimensional vectors under the model name candidate.
For a real model, the specification must name its endpoint, model, dimensions,
and document prefix. The proposed workflow uses an OpenAI-compatible endpoint
and a vector HNSW index with at most 2,000 dimensions.
prepare creates the candidate storage and bookkeeping, installs change
capture, and records where the run starts. backfill then computes the candidate
values. With PTAH_SPEC, PTAH_DB_URL, and PTAH_RUN_ID set as described in the
supporting files, these are ordinary native Ptah commands:
ptah inference prepareptah inference backfillAfter recovering from a simulated provider outage, the successful retry reports:
backfill finished: 5 scanned, 5 embedded, 0 skippedThose are the fixture’s rows, not a throughput measurement. The important result is that the candidate was filled while the old vectors remained available. The original payloads and index were preserved as well.
Catch up with the application
Section titled “Catch up with the application”The database does not stand still while a provider computes embeddings. Ptah’s outbox records relevant source changes in a companion table. PostgreSQL commits the source change and its event in the same transaction: if the write commits, its capture commits with it.
catchup reads captured changes and brings the candidate forward. An insert
needs a new embedding. A text update needs an embedding of the current text.
A delete is accounted for as a tombstone, a deletion marker that prevents a
late embedding result from restoring a removed source row.
The generated source column matters here too: changing the text inside the payload also changes the source value watched by the migration. The integration does not require an additional API call to tell Ptah that a memory changed.
After the first backfill, our test edits a memory, inserts another, and deletes one. It then runs catch-up. We check the replacement vector values for the edit and insert, and confirm that the deleted memory does not appear in search.
Once those changes are reconciled, build the candidate index and check readiness:
ptah inference catchupptah inference indexptah inference verifyCatch-up is a phase to repeat as needed, not a promise that future writes have already happened. Writers can create more work after a pass finishes. The final handover therefore includes another catch-up after writers have been paused.
Verification is about readiness, not better answers
Section titled “Verification is about readiness, not better answers”Verification checks whether source rows are accounted for, whether the candidate is current under the selected consistency mode, whether dimensions match, and whether its declared index is ready. The recorded result includes:
- every deterministic layer passedThat establishes migration readiness. It does not establish that the new model ranks relevant memories above irrelevant ones. A dimensionally correct vector can still represent the wrong text preprocessing, an unsuitable model, or a model that performs poorly for the application’s language and subject matter.
Keep the model and preprocessing choices explicit, then evaluate representative queries with the replacement model’s own query configuration. Our evaluation corpus reference describes how to record expected matches. Data coverage and retrieval quality answer different operational questions; neither substitutes for the other.
Cutover requires approval of a specific plan. Running it without approval refuses the operation and prints the plan digest:
ptah inference cutoverplan d2de8c26c114cutover refused: - this policy requires an approval and none was givenerror: cutover refusedThe digest identifies the plan the operator is authorizing. Review that plan,
then pass its digest to --approve with an approver name. If the evidence
changes, an earlier approval should not silently authorize a different switch.
Our fixture then approved that exact plan:
ptah inference cutover --approve d2de8c26c114 --approver "integration test"Use the digest printed by your own run; the value above belongs to this fixture.
Keep the application change small
Section titled “Keep the application change small”The application needs to select a prepared representation consistently. Our
Zengram change introduces PGVECTOR_COLUMN, defaulting to vector. The same
selection reaches distance calculations, ordinary upserts, and the transaction
that supersedes an older fact or status while inserting its replacement.
Startup checks that the selected column exists, is a pgvector vector column,
and has dimensions compatible with the encoder. The Ptah workflow owns candidate
index creation, so the API does not build a second index for it. The legacy
reembed.js refuses to run when a candidate column is selected, avoiding an
in-place rebuild of the original column under the new model configuration.
These changes stay in the storage layer and the legacy script’s guard. Search routes, keyword retrieval, entity extraction, and provider implementations keep their existing responsibilities. Ptah runs as an operator tool rather than an extra service inside the API deployment.
Coordinate the final handover
Section titled “Coordinate the final handover”Ptah’s cutover moves its active-generation pointer. Zengram does not read that pointer. Each API process selects its encoder and vector column from its configuration when it starts. Moving the pointer alone does not redirect a running API’s queries.
Pause all writers, including imports and consolidation jobs, and drain their
in-flight requests. Catch up again, verify, and approve the cutover. Deploy the
matching encoder configuration together with PGVECTOR_COLUMN=embedding_v2.
Check the new API’s search behavior and retire the old processes before writers
resume. Old processes can continue serving reads during preparation, but must
not resume writing under the old configuration after handover.
This preserves the old search path through candidate construction. It is not a zero-downtime deployment controller. Model rollout, database setup locks, provider capacity, and the final writer pause still need an operational plan. The dimension check also cannot identify two different models that happen to produce vectors of equal length.
After activation, schedule catch-up for the selected generation to keep its metadata current and drain captured changes. Preserve its specification and run identifier; retiring a generation while the API still reads its column would remove storage the application needs.
Before and after
Section titled “Before and after”| Area | Before Ptah | With this integration |
|---|---|---|
| Active storage | One vector column rebuilt in place | Original and candidate columns coexist |
| Search during backfill | Uses the representation being replaced | Uses the original encoder and vectors |
| Provider failure | Can leave gaps in the in-place rebuild | Candidate can be retried without clearing old vectors |
| Concurrent writes | Need coordination around the script’s row scan | Outbox captures changes for catch-up |
| New ANN index | Rebuilt on the active column | Built separately for the candidate |
| Readiness | Script results and operator checks | Explicit generation verification |
| Switch | Replacement of the existing representation | Approved generation plus coordinated API deployment |
| Rollback | Depends on rebuilds or backups | Old vectors remain initially, but become stale after new writes resume |
Keeping the old column is not a complete rollback mechanism. Before writes resume, it provides a way to cancel the application switch. Afterward, new and changed memories are written using the selected new encoder and column. The retained old vectors no longer describe the whole current corpus.
This integration does not register the legacy column as a Ptah generation and does not provide automatic rollback to it. Returning safely requires rebuilding or catching up that model’s representation and verifying it under a coordinated writer pause. Keep backups and both encoder configurations. The rollback and retirement guide explains the maintenance required when previous generations are managed by Ptah.
What the test establishes
Section titled “What the test establishes”The integration test runs the real Zengram storage code and released Ptah against PostgreSQL with pgvector. Its deterministic local embedding endpoint lets us simulate a provider failure without credentials or external model calls. It exercises the 1,536-to-384 dimension change, including a search through the old encoder while a candidate embedding request is still in flight.
The assertions cover retry without altering the original vectors or payloads, preservation of the original index, insert/update/delete catch-up, approval refusal and acceptance, and writes after the application selects the new column. Both ordinary upserts and fact/status supersession are exercised. Search also checks tenant and collection filtering, so switching columns does not discard those existing boundaries.
This is evidence for the migration mechanics on that fixture, not a production-scale latency benchmark or a claim about a model’s quality. The test source and transcript make that distinction inspectable.
An embedding model change gives stored data a new interpretation. Treat it as a versioned data migration: retain a usable representation, build its replacement, reconcile intervening writes, and verify readiness before the application selects it. In Zengram, that shift required an explicit column selection rather than a new search architecture.
Put Ptah to work
Example files
- candidate.json
- commands/approve.txt
- commands/build.txt
- commands/check.txt
- commands/cutover.txt
- file-hashes.json
- measured/approval-refusal.txt
- measured/backfill.txt
- measured/commands.json
- measured/transcript.txt
- measured/verify-excerpt.txt
- pgvector.js.txt
- README.md
- reembed.js.txt
- source.sql
- upstream-spec.json
- verified.json
- verify-article.js.txt
- verify.js.txt
- ZENGRAM-LICENSE.txt