Read one level up: https://search.qdrant.tech/md/documentation/tutorials-operations

> Explore Qdrant's agent skills catalog at https://skills.qdrant.tech/
> Search the documentation at https://skills.qdrant.tech/search?query=your+query+here
> Use this file to discover all available pages: https://qdrant.tech/llms.txt
# Migrate to a New Embedding Model with Zero Downtime in Qdrant

| Time: 40 min | Level: Intermediate |
| --- | ----------- |

When building a semantic search application, you need to [choose an embedding 
model](https://qdrant.tech/documentation/search-patterns/choose-embedding-model/index.md). Over time, you may want to switch to a different model for better
quality or cost-effectiveness. If your application is in production, this must be done with zero downtime to avoid 
disrupting users. Switching models requires re-embedding all vectors in your collection, which can take time.

This tutorial will guide you step-by-step through the two options for migrating to a new model with zero downtime, and shows how to [compare retrieval quality](?s=check-retrieval-quality-before-switching) on your own queries before you switch any traffic.

Re-embedding requires access to the original data used to create the embeddings. This data can come from a primary database, or it may be stored in the payloads of the points in Qdrant. This tutorial assumes that the necessary data is stored in the payloads. This is usually the case, as the payload often contains the text or other data that was used to generate the embeddings.

<aside role="status">
    Tested with Qdrant 1.19.0.
</aside>

The code examples in this tutorial use [Qdrant Cloud Inference](https://qdrant.tech/documentation/inference/cloud-inference/index.md) to generate vector embeddings. If you manage your own embedding infrastructure, you can apply the same principles, but you'll need to adapt the code examples for your embedding service.

## Two Options

The best approach to migrating to a new embedding model depends on how your collection has been configured. A blue-green migration (option 1) works with any collection type. Alternatively, if you use named vectors and your deployment is running version 1.18 or later, option 2 is easier, faster, and uses fewer resources.

### Option 1: Blue-Green Migration

The [blue-green migration approach](?s=blue-green-migration) uses two parallel collections. Start by creating a new collection configured for the new embedding model. Then, enable dual writes such that every incoming upsert is written to both collections simultaneously. Use a background scroll to re-embed each point using the new model, and write it to the new collection. Once migration is complete, switch search traffic to the new collection (flipping the alias, if applicable) and disable dual writes. This option works with any collection type, regardless of whether you use unnamed or named vectors. 

This approach has a couple of downsides:
- It duplicates payloads across both collections. For text-heavy collections where the payload is large, this can have a significant impact.
- Deletes or partial updates need to be paused during the migration or you need to implement additional logic to handle them.

### Option 2: Named Vectors

The [named vectors approach](?s=migrate-using-named-vectors) keeps everything in a single collection. Start by [adding the new model as an additional named vector](https://qdrant.tech/documentation/manage-data/collections/index.md#update-vector-schema): a schema-only operation that doesn't affect existing data. Next, enable dual writes so that every incoming upsert embeds with both models. Then, use a background scroll to update the new named vector on each existing point, leaving the old vector and payload intact. Once all points are re-embedded, you switch the `using` parameter in your search queries to the new vector, and then delete the old named vector.

The downside of this approach is that it only works for collections that were created with named vectors.

Compared to a blue-green migration, this approach:

- Doesn't require a second collection or any data copying.
- Keeps all point IDs, payloads, and other named vectors intact throughout the migration.
- Makes rollback trivial: the old named vector stays in the collection until you explicitly delete it.

When updating a point, make sure your dual-write logic also updates the new named vector at the same time. Updating only one will cause the two vectors to diverge. During [backfill](?s=step-3-re-embed-existing-points), searches and new inserts can continue, but updates and deletes to existing points must be paused.

## Blue-Green Migration

A blue-green migration uses two collections: the first collection contains the old embeddings, and the second one is used to store the new embeddings. A migration process copies the data from the old collection to the new one, re-embedding vectors using the new model. During the migration, you keep searching the old collection while writing any data updates to both collections. Once all vectors are re-embedded, switch the search to use the new collection.

![Blue-green embedding model migration: the application sends update events to the current embedding service and collection and, during the migration, also to a new embedding service and collection. A migration service reads the current collection, embeds with the new model, and inserts into the new collection only if the point does not exist yet.](/md/docs/embedding-model-migration.png)

*Blue-green embedding model migration. Select a step to see which services and collections are active; the bars are illustrative.*


The solution outlined here only works as-is for upsert operations. If you use deletes or partial updates, it is necessary to pause those operations during the migration or implement additional logic to handle them.

### Step 1: Create a New Collection

The first step is to create a new collection that will be used to store the new 
embeddings, compatible with the new model in terms of vector size and similarity function.

```python
client.create_collection(
    collection_name=NEW_COLLECTION,
    vectors_config=(
        models.VectorParams(
            size=512,  # Size of the new embedding vectors
            distance=models.Distance.COSINE  # Similarity function for the new model
        )
    )
)
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/create-new-collection/index.md).


Now is also a good moment to consider changing any other settings for the collection, like custom sharding, replication factor, etc. Switching the model may be a good opportunity to improve the performance of your search.

The newly created collection is empty and ready to be used for storing the new embeddings.

### Step 2: Enable Dual Writes

To ensure that both collections are kept up-to-date during the migration, write any changes to both collections simultaneously. This way, any new data or updates to existing data are reflected in both collections.

Ideally, the data in Qdrant is updated by an update service reading from an update queue. This service is responsible for embedding the documents and writing them to Qdrant. It uses code similar to this:

```python
client.upsert(
    collection_name=OLD_COLLECTION,
    points=[
        models.PointStruct(
            id=1,
            vector=models.Document(
                text="Example document",
                model=OLD_MODEL,
            ),
            payload={"text": "Example document"}
        )
    ]
)
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/upsert-old-collection/index.md).


To update the new collection, deploy a second service that updates the new collection in parallel with the existing one. This service uses the new embedding model to encode the documents and writes them to the new collection:

```python
client.upsert(
    collection_name=NEW_COLLECTION,
    points=[
        models.PointStruct(
            id=1,
            # Use the new embedding model to encode the document
            vector=models.Document(
                text="Example document",
                model=NEW_MODEL,
            ),
            payload={"text": "Example document"}
        )
    ]
)
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/upsert-new-collection/index.md).


A good practice is to always ensure that both operations succeed. Any errors need to be handled on the client side. You could store errors in a log or "dead letter queue" for later processing. Transient errors can be retried at a later time. Other errors need to be analyzed and addressed accordingly.

If you have a monolithic application instead of update services, you need to modify your application code to write to both collections simultaneously during the transition period. In your code, where you handle the embedding of the documents, you should add the logic to write to both collections.

Note that the method outlined in this tutorial only works for `upsert` operations. For example, a `delete` operation would fail on the new collection if a point does not exist yet, and that point would later be erroneously added by the migration process. If you use one of the following methods to modify points in your collection, you will need to pause those operations during the migration or implement additional logic to handle them:

- `.delete` - removing specified points from the collection
- `.update_vectors` - updating specified vectors on points
- `.delete_vectors` - deleting specified vectors from points
- `.set_payload` - setting payload values for specified points
- `.overwrite_payload` - overwriting the entire payload of a specified point with a new payload
- `.delete_payload` - deleting a specified key payload for points
- `.clear_payload` - removing the entire payload for specified points
- `.batch_update_points` - making batch updates to points, including their respective vectors and payloads

Refer to the [documentation of the SDK you are using](https://qdrant.tech/documentation/interfaces/index.md), or the 
[HTTP](https://api.qdrant.tech/api-reference)/[gRPC](https://api.qdrant.tech/api-reference) definitions, for the exact method names, as they may vary between languages.

After making these changes, you will be in a **dual-write mode**, where any change is written to both the old and new collection. This allows you to keep both collections up-to-date during the migration process.

### Step 3: Migrate the Existing Points into the New Collection

Now that you're in dual-write mode, it is time to migrate the existing points from the old collection to the new one. This can be done in a separate process that runs
in parallel with the regular upsert services. 

The migration process reads the points from the old collection, re-embeds them using the new model, and writes them to the new collection, making sure not to overwrite existing points inserted by the update service. Here's an example of what the code for such a migration process could look like:

```python
last_offset = None
batch_size = 100  # Number of points to read in each batch
reached_end = False

while not reached_end:
    # Get the next batch of points from the old collection
    records, last_offset = client.scroll(
        collection_name=OLD_COLLECTION,
        limit=batch_size,
        offset=last_offset,
        # Include payloads in the response, as we need them to re-embed the vectors
        with_payload=True,
        # We don't need the old vectors, so let's save on the bandwidth
        with_vectors=False,
    )

    # Re-embed the points using the new model
    points = [
        models.PointStruct(
            # Keep the original ID to ensure consistency
            id=record.id,
            # Use the new embedding model to encode the text from the payload,
            # assuming that was the original source of the embedding
            vector=models.Document(
                text=(record.payload or {}).get("text", ""),
                model=NEW_MODEL,
            ),
            # Keep the original payload
            payload=record.payload
        )
        for record in records
    ]

    # Upsert the re-embedded points into the new collection
    client.upsert(
        collection_name=NEW_COLLECTION,
        points=points,
        # Only insert the point if a point with this ID does not already exist.
        update_mode=models.UpdateMode.INSERT_ONLY
    )

    # Check if we reached the end of the collection
    reached_end = (last_offset == None)
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/migrate-points/index.md).


Breaking down this code step by step:

- Data is read from the old collection in batches of 100 points using a [scroll](https://qdrant.tech/documentation/manage-data/points/index.md#scroll-points).
- For each batch of points, the process re-embeds the vectors using the new embedding model. It assumes that the original text used for embedding is stored in the payload under the key `text`.
- With the re-embedded vectors, it upserts the points into the new collection, keeping the original IDs and payloads. The upserts use [insert-only mode](https://qdrant.tech/documentation/manage-data/points/index.md#update-mode) to ensure that a point is only inserted if it does not already exist in the new collection (available in version 1.16 or later). This prevents overwriting newer updates from the regular update service.

The migration process can take some time, and the offset can be stored in a persistent way so you can resume the migration process in case of a failure. You can use a database, a file, or any other persistent storage to keep track of the last offset. Having said that, because the conditional upserts would not overwrite any points in the new collection, you could safely restart the migration process from the beginning if needed.

### Step 4: Change the Collection and Embedding Model for Searches

Once the migration process is complete and all the points from the old collection are re-embedded and stored in the new collection, [compare retrieval quality](?s=check-retrieval-quality-before-switching) between the two collections on your own queries. If you have not deleted any points since the migration started, `count` on both collections returns the same number, which is a quick completeness check. When the new collection performs at least as well, you can roll out a configuration change of the backend application. There are two key changes you have to make:

1. **The collection name**. Switch this from the old collection to the new collection. If you're using a [collection alias](https://qdrant.tech/documentation/manage-data/collections/index.md#collection-aliases), switch the alias to point to the new collection. Deleting and recreating the alias in a single request makes the switch atomic:

```python
client.update_collection_aliases(
    change_aliases_operations=[
        models.DeleteAliasOperation(
            delete_alias=models.DeleteAlias(alias_name="prod")
        ),
        models.CreateAliasOperation(
            create_alias=models.CreateAlias(
                collection_name=NEW_COLLECTION, alias_name="prod"
            )
        ),
    ]
)
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/alias-update/index.md).


2. **The embedding model**. Switch this from the old embedding model to the new embedding model.

If these values are hardcoded in your application, you will need to change them directly in the code and deploy a new version of your application. For example, if your current search code looks like this:

```python
results = client.query_points(
    collection_name=OLD_COLLECTION,
    query=models.Document(text="my query", model=OLD_MODEL),
    limit=10,
)
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/search-old-collection/index.md).


You need to change it in the following way:

```python
results = client.query_points(
    collection_name=NEW_COLLECTION,
    query=models.Document(text="my query", model=NEW_MODEL),
    limit=10,
)
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/search-new-collection/index.md).


### Step 5: Wrapping Up

Once your application has switched to the new collection, keep the dual-write mode from Step 2 running for an observation period. While the old collection still receives every write, you can roll back at any time by pointing the alias (or the collection name) back at the old collection and switching the embedding model back. Run the same call as in Step 4 with `OLD_COLLECTION`.

When you are confident in the new model, disable dual writes. From now on, the application should only write to the new collection. Rolling back after this point means losing every write since, because the old collection no longer receives updates.

All searches are now performed using the new embeddings. If the old collection is no longer needed, you can safely delete it. To keep the option of restoring it, take a snapshot of the old collection first.

---

## Migrate Using Named Vectors

If your collection uses [named vectors](https://qdrant.tech/documentation/manage-data/points/#named-vectors/index.md) and your deployment is running version 1.18 or later, you can migrate to a new embedding model without creating a second collection. Instead, [add the new model as an additional named vector to the existing collection's schema](https://qdrant.tech/documentation/manage-data/collections/index.md#update-vector-schema), re-embed points in the background, switch the `using` parameter in your search queries, and then delete the old named vector.

This approach only works when your collection was created with named vectors and your deployment is running version 1.18 or later. If not, use a [blue-green migration](?s=blue-green-migration) instead.

### Step 1: Add the New Named Vector

Add the new model's vector schema to the existing collection. This is a schema-only operation: no segments are rebuilt and no existing point data is modified. The new vector is queryable immediately, but queries return no results until points are populated with values for it.

```python
client.create_vector_name(
    collection_name=COLLECTION,
    vector_name=NEW_VECTOR,
    vector_name_config=models.DenseVectorNameConfig(
        dense=models.DenseVectorConfig(
            size=512,  # Size of the new embedding vectors
            distance=models.Distance.COSINE  # Similarity function for the new model
        )
    ),
)
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/add-named-vector/index.md).


### Step 2: Enable Dual Writes

Update your upsert service to embed each document with both models and write both named vectors on every upsert:

```python
client.upsert(
    collection_name=COLLECTION,
    points=[
        models.PointStruct(
            id=1,
            vector={
                OLD_VECTOR: models.Document(
                    text="Example document",
                    model=OLD_MODEL,
                ),
                NEW_VECTOR: models.Document(
                    text="Example document",
                    model=NEW_MODEL,
                ),
            },
            payload={"text": "Example document"}
        )
    ]
)
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/upsert-both-vectors/index.md).


From this point on, every new or updated point carries both embeddings.

### Step 3: Re-Embed Existing Points

<aside role="status">
Searches and new inserts can continue during backfill. Pause updates and deletes to existing points, and wait for those in-flight operations to finish before starting. A concurrent update could change the payload after it is read, causing the backfill to overwrite the new vector with an embedding of stale text. A concurrent delete could cause the batch to fail.
</aside>

Run a background process that scrolls through the collection and updates only the new named vector on each existing point. Because `update_vectors` is used rather than `upsert`, the old named vector and the payload on each point remain unchanged.

```python
last_offset = None
batch_size = 100
reached_end = False

while not reached_end:
    records, last_offset = client.scroll(
        collection_name=COLLECTION,
        limit=batch_size,
        offset=last_offset,
        with_payload=True,
        with_vectors=False,
    )

    # Update only the new vector on each point; the old vector and payload are untouched
    client.update_vectors(
        collection_name=COLLECTION,
        points=[
            models.PointVectors(
                id=record.id,
                vector={
                    NEW_VECTOR: models.Document(
                        text=(record.payload or {}).get("text", ""),
                        model=NEW_MODEL,
                    )
                },
            )
            for record in records
        ],
    )

    reached_end = last_offset is None
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/re-embed-existing/index.md).


Once the backfill has completed successfully, resume updates and deletes with dual writes enabled so every new or updated point continues to receive both embeddings.

### Step 4: Switch Search to the New Vector

Before switching, check that no point is missing the new vector, then [compare retrieval quality](?s=check-retrieval-quality-before-switching) between the two vectors on your own queries. This count returns `0` when the re-embedding is complete:

```python
missing = client.count(
    collection_name=COLLECTION,
    count_filter=models.Filter(
        must_not=[models.HasVectorCondition(has_vector=NEW_VECTOR)]
    ),
    exact=True,
).count
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/count-missing/index.md).


When every point has the new vector and it performs at least as well, change the query logic:
- switch the `using` parameter from the old vector to the new vector.
- switch the embedding model from the old model to the new model. 

Before:

```python
results = client.query_points(
    collection_name=COLLECTION,
    query=models.Document(text="my query", model=OLD_MODEL),
    using=OLD_VECTOR,
    limit=10,
)
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/search-with-old-vector/index.md).


After:

```python
results = client.query_points(
    collection_name=COLLECTION,
    query=models.Document(text="my query", model=NEW_MODEL),
    using=NEW_VECTOR,
    limit=10,
)
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/search-with-new-vector/index.md).


### Step 5: Disable Dual Writes and Delete the Old Named Vector

Once all search traffic uses the new vector, keep dual writes running for an observation period. Until you delete the old vector, you can roll back by switching the `using` parameter and the embedding model back to the old values.

When you are confident in the new model, change your upsert service to write only to the new vector going forward. Next, delete the old named vector from the collection. Deleting it is the point of no return: queries that use the old vector name fail afterwards, and the old embeddings can only be restored by re-embedding.

```python
client.delete_vector_name(
    collection_name=COLLECTION,
    vector_name=OLD_VECTOR,
)
```

> This snippet is also available in Typescript, Rust, Java, Csharp, and Go. See the [full snippet](https://qdrant.tech/documentation/snippets/tutorial-model-migration/delete-old-named-vector/index.md).


The old vector's storage is reclaimed after the next optimizer run. All point IDs, payloads, and the new named vector remain intact.

## Check Retrieval Quality Before Switching

A new embedding model is not automatically better on your data, so measure it before you move traffic. Both embeddings exist side by side at this point (two collections in a blue-green migration, two named vectors in the other option), which makes it possible to run identical queries against each and compare the results.

Start from a labeled set of queries, with the documents that should be returned for each one. [Measuring Retrieval Relevance](https://qdrant.tech/documentation/search-evaluation/retrieval-relevance/index.md) explains how to build it from logs, human annotation, or synthetic queries, and how to choose metrics. Use real queries from your application where you can, because a model that wins on generic benchmarks can still lose on your domain.

To score both setups, follow these steps:

1. **Query each setup with the same inputs.** For every query in the labeled set, ask the old setup and the new setup for the top `k` results. Use the same `k` for both. In a blue-green migration, that means querying the old and new collections, each with its own model. With named vectors, query the one collection twice and select the old or new vector each time, again with the matching model.
2. **Collect the ranked results per query.** Record the document IDs in the order returned, along with their scores. Make sure the IDs use the same format as the IDs in your labeled set (for example, strings on both sides), or nothing will match.
3. **Compute the metrics against the labels.** Use an evaluation library such as [ranx](https://amenra.github.io/ranx/) for Python, or implement the metrics yourself in other languages. Three metrics cover most cases:
   - **Recall@10:** the share of the relevant documents that appear in the top 10.
   - **MRR:** the average of 1 divided by the rank of the first relevant result, which rewards putting a correct answer near the top.
   - **nDCG@10:** a score that rewards relevant documents appearing higher in the list, and handles graded relevance.
4. **Compare the two score sets.**

Read the comparison with these points in mind:

- **Decide the bar first.** Pick what metric to look at, and how big the gains have to be to justify the switch to a new embedding model. Different metrics can rank the two models differently.
- **Look at individual queries.** Average can often hide a group of queries that got worse on the new embedding model. Sort the query by score and read the biggest regression.
- **Weigh the costs.** A model with larger vectors needs more memory and disk, and may change search latency. Compare these on the new collection as well as the quality scores.
- **Without labels, you can only measure difference.** The overlap between the top results of both models tells you how much the switch changes what users see, not which results are better. Review the queries with the lowest overlap by hand.

If the new model does not clear the bar, stop here. Production still serves the old embeddings, so the only cleanup is deleting the new collection, or the new named vector with `delete_vector_name`.