Skip to content

Changing Embedding Models: Zengram Before and After Ptah

We tested a small Zengram integration that builds replacement embeddings with Ptah while old search keeps serving, then coordinates the final handover.

Ptah helps you plan, review, and apply database migrations. Try in your browser

A search request arrives while an application is changing its embedding model. The new model produces 384 numbers per document. The database still contains 1,536 numbers per document, calculated by the previous model. Changing an API setting cannot make those stored values belong to the new model.

For a small corpus, rebuilding everything during maintenance may be enough. For a corpus with millions of rows, that rebuild could become the longest part of the deployment. The search service needs a usable representation throughout that work, including when the replacement provider fails.

Zengram is a shared memory service for AI agents. It stores memories in PostgreSQL and retrieves them through vector, keyword, and entity-based search. Its embedding replacement script gave us a concrete place to test a different migration shape: build a new representation beside the active one, check it, and switch deliberately.

We implemented and tested that approach in an upstream proposal. The PR is open; this is a case study of our tested integration, not an announcement of Zengram adopting Ptah. The examples use Ptah 0.11.4, PostgreSQL 16.15, and pgvector 0.8.7. The supporting files contain the specification, test inputs, and recorded output.

An embedding is a list of numbers that a model calculates from an input such as a sentence. Similar inputs should occupy nearby positions in that model’s vector space. A vector’s dimension is the length of the list: 1,536 values or 384 values in this example.

To search, the application embeds the query and compares that vector with stored document vectors. Both sides must use compatible coordinates. Imagine looking for “database connection failures” among stored troubleshooting notes: the query and each note need to be represented by the same model, with the appropriate query and document preprocessing.

Stored text and a search query pass through encoder A before comparison. A query sent through encoder B reaches an incompatible vector space.

Different dimensions make a mismatch obvious: PostgreSQL cannot compare the vectors as if their lengths matched. Equal dimensions do not make different models interchangeable. Two models can each return 384 values while assigning those coordinates different meanings. A comparison can then execute and still produce meaningless rankings.

PostgreSQL’s pgvector extension supplies vector storage and distance operators. An HNSW index builds a graph that helps find nearby vectors without comparing the query with every row. That graph belongs to the vectors it indexes. Changing the model therefore coordinates stored data, the search index, and the encoder configuration used by the application.

Before Ptah: rebuild the representation being searched

Section titled “Before Ptah: rebuild the representation being searched”

Zengram’s api/scripts/reembed.js reads each memory’s persisted text and asks the currently configured provider to embed it. Its destination is the same memories.vector column that the API searches.

When the script detects a dimension change, it drops idx_memories_vector_hnsw, changes the column to the new vector dimension using NULL values, and then writes replacement embeddings row by row. At the end it rebuilds HNSW. Payloads and the separate keyword-search data remain in place.

The API searches memories.vector while the dimension-change script drops its index, clears the vectors, fills the same column, and rebuilds HNSW.

The script has useful operational controls: dry-run is the default, embedding calls have bounded concurrency, and a failed row is reported. It can be rerun. Those controls do not preserve the old vector representation once the dimension change has cleared it. A provider failure can leave some rows without vectors until a later attempt succeeds.

This is a manageable procedure when maintenance is short and its search impact is acceptable. As the corpus grows, completion time and provider availability become conditions for restoring the active representation. Zengram’s keyword fallback can still help answer requests, but it is not the old vector search.

A new column alone is not a finished migration

Section titled “A new column alone is not a finished migration”

Building replacement vectors takes more than adding vector(384) to a table. The provider can fail halfway through. A backfill—the pass that computes values for existing rows—can take much longer than expected. A completed backfill can still be missing a newly inserted memory, or contain the earlier text of a memory edited while the provider was processing it.

Deletes matter too. An embedding request already in flight must not bring a deleted memory back into the candidate representation. The new index must be ready before it is relied upon. Document prefixes, truncation rules, and model selection must agree with the configuration that will produce future writes.

Finally, deploying the new query encoder before its matching stored vectors are ready creates a mismatch at the application boundary. This is a migration of data, model, index, and application configuration. The table alteration is only its starting point.

A generation groups vectors produced with a particular model, preprocessing configuration, and destination. In our integration, the original vectors stay in memories.vector. The candidate goes into memories.embedding_v2.

The serving path uses the API, old encoder, original vector column, and old HNSW. A separate Ptah path backfills embedding_v2, catches up changes, builds its index, and verifies it.

The API keeps its existing encoder during this build. Ptah uses the replacement encoder to populate the candidate. A failed candidate request does not clear the original column or drop its index. It changes whether the replacement is ready, rather than removing the representation that already works.

Use a new column and run identifier for each model change, including changes that preserve dimensions. This keeps a retry of the same build distinct from starting a different generation.

Expose the text the application actually stores

Section titled “Expose the text the application actually stores”

Zengram keeps source text inside a JSON payload. Ptah’s source specification reads columns. The integration adds a stored generated embedding_text column that follows the same text fallback as reembed.js: text, then content, then note, then title.

examples/source.sql
-- Expose the same text fallback as reembed.js to Ptah's column-based source.
-- A generated column follows payload edits without another application write.
-- Adding a STORED column rewrites the table: schedule this one-time setup.
SET lock_timeout = '5s';
ALTER TABLE public.memories ADD COLUMN embedding_text text GENERATED ALWAYS AS (
COALESCE(NULLIF(payload->>'text', ''), NULLIF(payload->>'content', ''),
NULLIF(payload->>'note', ''), payload->>'title', '')
) STORED;

The generated expression follows payload edits automatically. It avoids asking every application writer to maintain another copy of the text. This has a real setup cost: adding a stored generated column rewrites and locks the table. Schedule that one-time operation separately. The five-second lock timeout bounds waiting to acquire a lock, not the duration of the rewrite after acquisition.

The complete specification selects public.memories, uses id as its row key, reads embedding_text, and targets embedding_v2. It declares a mutable source with outbox consistency, a cosine-distance HNSW index, and a policy requiring explicit cutover approval.

The replacement endpoint in this worked example is a deterministic local test provider. It returns 384-dimensional vectors under the model name candidate. For a real model, the specification must name its endpoint, model, dimensions, and document prefix. The proposed workflow uses an OpenAI-compatible endpoint and a vector HNSW index with at most 2,000 dimensions.

prepare creates the candidate storage and bookkeeping, installs change capture, and records where the run starts. backfill then computes the candidate values. With PTAH_SPEC, PTAH_DB_URL, and PTAH_RUN_ID set as described in the supporting files, these are ordinary native Ptah commands:

examples/commands/build.txt
ptah inference prepare
ptah inference backfill

After recovering from a simulated provider outage, the successful retry reports:

examples/measured/backfill.txt
backfill finished: 5 scanned, 5 embedded, 0 skipped

Those are the fixture’s rows, not a throughput measurement. The important result is that the candidate was filled while the old vectors remained available. The original payloads and index were preserved as well.

The database does not stand still while a provider computes embeddings. Ptah’s outbox records relevant source changes in a companion table. PostgreSQL commits the source change and its event in the same transaction: if the write commits, its capture commits with it.

Ptah installs the outbox before backfill. The live API commits a memory change together with its outbox event. Catch-up refreshes the candidate or records a deletion while old search remains active.

catchup reads captured changes and brings the candidate forward. An insert needs a new embedding. A text update needs an embedding of the current text. A delete is accounted for as a tombstone, a deletion marker that prevents a late embedding result from restoring a removed source row.

The generated source column matters here too: changing the text inside the payload also changes the source value watched by the migration. The integration does not require an additional API call to tell Ptah that a memory changed.

After the first backfill, our test edits a memory, inserts another, and deletes one. It then runs catch-up. We check the replacement vector values for the edit and insert, and confirm that the deleted memory does not appear in search.

Once those changes are reconciled, build the candidate index and check readiness:

examples/commands/check.txt
ptah inference catchup
ptah inference index
ptah inference verify

Catch-up is a phase to repeat as needed, not a promise that future writes have already happened. Writers can create more work after a pass finishes. The final handover therefore includes another catch-up after writers have been paused.

Verification is about readiness, not better answers

Section titled “Verification is about readiness, not better answers”

Verification checks whether source rows are accounted for, whether the candidate is current under the selected consistency mode, whether dimensions match, and whether its declared index is ready. The recorded result includes:

examples/measured/verify-excerpt.txt
- every deterministic layer passed

That establishes migration readiness. It does not establish that the new model ranks relevant memories above irrelevant ones. A dimensionally correct vector can still represent the wrong text preprocessing, an unsuitable model, or a model that performs poorly for the application’s language and subject matter.

Keep the model and preprocessing choices explicit, then evaluate representative queries with the replacement model’s own query configuration. Our evaluation corpus reference describes how to record expected matches. Data coverage and retrieval quality answer different operational questions; neither substitutes for the other.

Cutover requires approval of a specific plan. Running it without approval refuses the operation and prints the plan digest:

examples/commands/cutover.txt
ptah inference cutover
examples/measured/approval-refusal.txt
plan d2de8c26c114
cutover refused:
- this policy requires an approval and none was given
error: cutover refused

The digest identifies the plan the operator is authorizing. Review that plan, then pass its digest to --approve with an approver name. If the evidence changes, an earlier approval should not silently authorize a different switch. Our fixture then approved that exact plan:

examples/commands/approve.txt
ptah inference cutover --approve d2de8c26c114 --approver "integration test"

Use the digest printed by your own run; the value above belongs to this fixture.

The application needs to select a prepared representation consistently. Our Zengram change introduces PGVECTOR_COLUMN, defaulting to vector. The same selection reaches distance calculations, ordinary upserts, and the transaction that supersedes an older fact or status while inserting its replacement.

Startup checks that the selected column exists, is a pgvector vector column, and has dimensions compatible with the encoder. The Ptah workflow owns candidate index creation, so the API does not build a second index for it. The legacy reembed.js refuses to run when a candidate column is selected, avoiding an in-place rebuild of the original column under the new model configuration.

These changes stay in the storage layer and the legacy script’s guard. Search routes, keyword retrieval, entity extraction, and provider implementations keep their existing responsibilities. Ptah runs as an operator tool rather than an extra service inside the API deployment.

Ptah’s cutover moves its active-generation pointer. Zengram does not read that pointer. Each API process selects its encoder and vector column from its configuration when it starts. Moving the pointer alone does not redirect a running API’s queries.

Pause and drain writers, run final catch-up, verify, approve cutover, deploy the new encoder and column together, check search and retire old API processes, then resume writers.

Pause all writers, including imports and consolidation jobs, and drain their in-flight requests. Catch up again, verify, and approve the cutover. Deploy the matching encoder configuration together with PGVECTOR_COLUMN=embedding_v2. Check the new API’s search behavior and retire the old processes before writers resume. Old processes can continue serving reads during preparation, but must not resume writing under the old configuration after handover.

This preserves the old search path through candidate construction. It is not a zero-downtime deployment controller. Model rollout, database setup locks, provider capacity, and the final writer pause still need an operational plan. The dimension check also cannot identify two different models that happen to produce vectors of equal length.

After activation, schedule catch-up for the selected generation to keep its metadata current and drain captured changes. Preserve its specification and run identifier; retiring a generation while the API still reads its column would remove storage the application needs.

Area Before Ptah With this integration
Active storage One vector column rebuilt in place Original and candidate columns coexist
Search during backfill Uses the representation being replaced Uses the original encoder and vectors
Provider failure Can leave gaps in the in-place rebuild Candidate can be retried without clearing old vectors
Concurrent writes Need coordination around the script’s row scan Outbox captures changes for catch-up
New ANN index Rebuilt on the active column Built separately for the candidate
Readiness Script results and operator checks Explicit generation verification
Switch Replacement of the existing representation Approved generation plus coordinated API deployment
Rollback Depends on rebuilds or backups Old vectors remain initially, but become stale after new writes resume

Keeping the old column is not a complete rollback mechanism. Before writes resume, it provides a way to cancel the application switch. Afterward, new and changed memories are written using the selected new encoder and column. The retained old vectors no longer describe the whole current corpus.

This integration does not register the legacy column as a Ptah generation and does not provide automatic rollback to it. Returning safely requires rebuilding or catching up that model’s representation and verifying it under a coordinated writer pause. Keep backups and both encoder configurations. The rollback and retirement guide explains the maintenance required when previous generations are managed by Ptah.

The integration test runs the real Zengram storage code and released Ptah against PostgreSQL with pgvector. Its deterministic local embedding endpoint lets us simulate a provider failure without credentials or external model calls. It exercises the 1,536-to-384 dimension change, including a search through the old encoder while a candidate embedding request is still in flight.

The assertions cover retry without altering the original vectors or payloads, preservation of the original index, insert/update/delete catch-up, approval refusal and acceptance, and writes after the application selects the new column. Both ordinary upserts and fact/status supersession are exercised. Search also checks tenant and collection filtering, so switching columns does not discard those existing boundaries.

This is evidence for the migration mechanics on that fixture, not a production-scale latency benchmark or a claim about a model’s quality. The test source and transcript make that distinction inspectable.

An embedding model change gives stored data a new interpretation. Treat it as a versioned data migration: retain a usable representation, build its replacement, reconcile intervening writes, and verify readiness before the application selects it. In Zengram, that shift required an explicit column selection rather than a new search architecture.

Put Ptah to work

Example files