Upgrading Uptrace
Every Uptrace release can change the PostgreSQL and ClickHouse schemas. The uptrace migrate command applies those changes and moves the running processes to the new version, so Uptrace keeps accepting telemetry during the upgrade.
The short version for a DEB or RPM install:
# 1. Back up PostgreSQL and ClickHouse.
# 2. Install the new package (do not restart the service yet).
# 3. Rewrite deprecated config options, then check the config.
uptrace config fix
uptrace config fix --write
uptrace config validate
# 4. Preview the upgrade.
uptrace migrate up --plan
# 5. Apply the migrations and restart Uptrace.
uptrace migrate up --restart
The rest of this page explains each step, what to do when a step fails, and how to roll back.
Supported upgrade paths
You can upgrade from each of these releases directly to the latest release, in one step. You do not need to install the releases between them:
| You run | Upgrade |
|---|---|
| v2.1.0 or later | Follow this page. |
| v2.0.0-beta.3 to v2.0.3, v2.1.0-beta to v2.1.0-beta.8 | Read Upgrading from v2.0 or a v2.1 beta first, then follow this page. |
| v2.0.0-beta.1, v2.0.0-beta.2, v2.0.0-rc.2 | Not supported. Install the latest release on new databases. |
| v1.x | Not supported. v2 stores data differently. Run v2 next to v1 as the v2.0 migration strategy describes, then remove v1. |
Every release in the table is tested: its own binary creates the databases, and the latest release then upgrades them.
Before you upgrade
Before you start:
- Read the release notes on GitHub Releases.
- Back up both PostgreSQL and ClickHouse. The Docker, Kubernetes, and Ansible guides show how.
- With the new binary installed, check that it can start against your stores:shell
uptrace preflightpreflightchecks that the binary has the license keys of its edition, that PostgreSQL, every ClickHouse replica, and the Kafka brokers (when configured) answer with your credentials, that PostgreSQL and ClickHouse run supported versions, and that no migration stopped partway and no rollback was left unfinished. It exits non-zero when one of these fails. Unapplied migrations are expected at this point, so they only warn.preflightruns a subset ofuptrace doctor, which also checks Redis, the replicas, the disks, and the outbound endpoints.
All uptrace migrate commands read /etc/uptrace/config.yml by default. Pass --config=/path/to/config.yml to use another file.
Upgrading from v2.0 or a v2.1 beta
v2.0.x and the v2.1.0 betas are classic releases: they migrated the databases with other commands and recorded the migrations in other tables. The latest release upgrades their databases, but the first upgrade needs these extra steps.
- Upgrade the databases. Uptrace now needs ClickHouse 26.3 or later and PostgreSQL 15 or later. Upgrade the servers before you install the new Uptrace: v2.0.x and the v2.1 betas keep working on ClickHouse 26.3.
uptrace preflightchecks both versions. - Set a secret. The new release refuses to start when
service.secretis the placeholderFIXMEthat older configs shipped with:textinvalid config: service.secret is the shipped placeholder "FIXME"; set a random value
Set it to a random value:shellopenssl rand -hex 32
The secret signs login sessions, so a new secret logs out every user once. The DEB and RPM packages replace the placeholder for you. - Rewrite the config. Run
uptrace config fix --write. It moves the options that changed place, for examplespanstopipelines.spansandch_schema.metricstoch_schema.timeseriesandch_schema.datapoints, and deletes the options that the new release does not use. - Replace the old migration commands in your scripts, init containers, and entrypoints.
serveno longer migrates the databases when it starts, as v2.0 did:Old command New command uptrace pg init,uptrace pg migrateuptrace migrate upuptrace ch init,uptrace ch migrateuptrace migrate upuptrace ch status,uptrace ch checkuptrace migrate statusuptrace sync_dashboardsuptrace dashboard syncuptrace migrate upmigrates PostgreSQL and ClickHouse together. The old commands print the new command and exit with an error. - Restart the old processes yourself. A classic process cannot receive the restart request of
--restart. After the expand steps,uptrace migrate upchecks PostgreSQL for the sessions of a classic process. While one is connected,upwaits up to--wait, then stops before the backfill, names the sessions, and exits non-zero. The old release keeps serving. Use the manual restarts case for the first upgrade:shelluptrace migrate up --before-release-change sudo systemctl restart uptrace uptrace migrate up
Later upgrades can use--restart.
On the first run, migrate reads the records of the classic release and does not run those migrations again. Before that run, status and preflight count every migration as unapplied. up --plan tells you how many steps it will take over:
The classic runners applied 7 steps that have no step record. They read as applied, and the next up, mark_applied, mark_unapplied, or down records them.
up then prints Adopted 7 steps that the classic runners applied. and runs the rest. It leaves the old bun_migrations and ch_migrations tables in place.
The DEB and RPM packages do not restart Uptrace onto databases that still need migrations. The old process keeps serving until you run the commands above.
How migrations run
A migration can have up to three phases, and the processes switch to the new version in the middle:
old version serves traffic
│ expand add new tables, columns, and indexes; the old version keeps working
▼
restart: every Uptrace process moves to the new version
hold: some migrations pause a capability, for example alerts, while the next phases run
│ backfill copy and convert existing rows into the new shape
│ finalize remove what only the old version needed; build the final indexes
▼
release: the paused capabilities start again
- Expand runs while the old version still serves, so it only adds objects the old version can ignore.
- Restart required is the point where every
serveandworkerprocess must run the new binary before the work continues. - Hold and release pause and resume a capability for the duration of the backfill and finalize phases. No process restarts for them.
- Backfill and finalize run on the new version. A long backfill commits in batches, so a stop does not lose its progress.
uptrace migrate records every completed step in PostgreSQL. A second run skips completed steps and continues where the previous run stopped.
Preview the upgrade
Use three read-only commands before you change anything. None of them takes a lock or writes to the databases.
See what the databases hold now:
uptrace migrate status
Status behind
Applied 41
Unapplied 2
20260915101500_heatmap unapplied
20260916090000_retention unapplied
Run: uptrace migrate up --plan
See what up will do, in order:
uptrace migrate up --plan
Plan 2 steps of 1 migration
Migrations
20260922080949_demo_split_customer_name split the customer name into first and last
# migration step runs as
- restart required every process must run the new release
- hold every process stops alerts and acks
1 20260922080949_demo_split_customer_name 2-backfill.up.pg.go 1 transaction per batch, resumes with cursor=4
2 20260922080949_demo_split_customer_name 3-finalize.up.pg.sql 6 statements, no transaction; a stop resumes the file
- release every process starts alerts again
Completed, not run again
migration step completed at
20260922080949_demo_split_customer_name 1-expand.up.pg.sql 2026-09-22 08:14:03 UTC, group 2
Transactions 1 per batch of step 1
Restart required once: before step 1
Holds alerts: held before step 1, released after step 2
In this example, an earlier run already completed the expand step, so the plan continues the backfill at its last batch. The Restart required line tells you whether the processes must switch versions, and the Holds line tells you which capabilities pause and for how long. Add --plan to the exact command you intend to run: the same flags select the same work.
To read the SQL itself, print it without running it:
uptrace migrate show
uptrace migrate show --database pg
uptrace migrate show --database ch
Choose how the processes restart
Every upgrade runs in the same order:
- The expand steps run while the old version serves.
- Every
serveandworkerprocess moves to the new version. - The backfill and finalize steps run on the new version.
What differs is who moves the processes in step 2:
| Case | Who restarts the processes | Commands | Use it with |
|---|---|---|---|
| Built-in restarts | uptrace migrate | up --restart | systemd, Ansible |
| Manual restarts | You | up --before-release-change, then restart, then up | systemd, any supervisor |
| Kubernetes | Your deploy tool | The same as manual restarts, run as Jobs | Helm, Argo CD |
Manual restarts and Kubernetes are the same case: up stops before step 2, something else changes the version, and a second up finishes the work. On Kubernetes, the deploy tool changes the image, and the two up runs are Jobs.
Use the same steps for a release that adds no migration. Then up runs nothing, and only the processes restart.
Case 1: built-in restarts
uptrace migrate up --restart does the whole upgrade in one run. It needs a supervisor that starts a stopped process again from the installed binary. The packaged systemd unit sets Restart=always, so this works for the DEB, RPM, and binary installs. The Ansible playbook uses this case.
- Install the new version using DEB, RPM, or a pre-compiled binary. Do not restart the service.
- Fix and validate the config. A new release can move or rename config options. The old keys keep working, but Uptrace logs a deprecation warning for each one. Review the fixes, apply them, and then validate the result:shell
uptrace config fix uptrace config fix --write uptrace config validateconfig fixwithout--writeprints each fix and a diff, and changes nothing.--writerewrites the file and keeps the original asconfig.yml.bak. See Fixing deprecated options. - Apply the migrations and restart Uptrace:shell
uptrace migrate up --restart
up --restart runs the expand steps while the old version serves, then asks each running process to stop. systemd starts it again from the new binary. Workers restart first, then each serve process one at a time, so the API keeps serving when you run more than one. After that, up runs the backfill and finalize steps. At the end, it restarts every process that the run did not restart yet, so each process reads the new binary, config, and data files.
Use the same command after you change only the config or a data file: it restarts every process even when no migration is left.
--wait bounds how long up waits for each restart (2 minutes by default):
uptrace migrate up --restart --wait 5m
If a process does not come back in time, up stops and names it. Fix the cause, for example a host where the new binary is not installed, and run the same command again.
Case 2: manual restarts
Use this case when you restart the processes yourself, for example with systemctl on each host, or with a supervisor that --restart cannot use.
- Install the new version without a restart. Run
uptrace config fix --writeto rewrite deprecated options, and runuptrace config validate. - Run the expand steps.
upstops before the restart and exits zero. The old version keeps serving:shelluptrace migrate up --before-release-change - Restart every Uptrace process on every host, so it runs the new binary:shell
sudo systemctl restart uptrace - Run the backfill and finalize steps:shell
uptrace migrate up
Before the backfill, step 4 checks that every live process runs the new version. If one still runs the old version, up stops and names it. Restart that process and run up again, or pass --wait 10m to let up wait for the restarts:
uptrace migrate up --wait 10m
Case 3: Kubernetes
Kubernetes is case 2, with the deploy tool in step 3. The kubelet starts a stopped container again from the image of its pod spec, so only the deploy tool can move a pod to the new version. Run each up as a Job with the new image:
1. pre-upgrade Job uptrace migrate up --before-release-change
2. deploy tool change the image, and wait until the new pods are ready
3. post-upgrade Job uptrace migrate up --wait 10m
- Do not use
--restarton Kubernetes. It would only restart each pod again on the same image. - Start step 3 only after the rollout completes. With Helm, run
helm upgrade --wait, so Helm runs the post-upgrade hooks after the resources are ready. With Argo CD, use a PostSync hook. If a pod still runs the old image, step 3 waits up to--waitfor it, then stops and names the pod. Run the Job again after the rollout.
The Helm chart runs these steps for you. A pre-upgrade Job runs step 1 with the new image and the new config, helm upgrade changes the image in step 2, and a post-upgrade Job runs step 3. If a Job fails, the upgrade stops, and the old pods keep serving. See upgrading with Helm.
Docker
Pull the new images and recreate the containers as described in upgrading with Docker.
Check the result
When the upgrade completes, status reports that the databases are up to date:
uptrace migrate status
Status up to date
Applied 43
In a script or a health check, use --fail-on-unapplied. It exits non-zero when a migration is unapplied or a rollback did not finish:
uptrace migrate status --fail-on-unapplied
To check what the upgraded processes see, read the report of each running process. It fails when a process serves with migrations that are not applied:
uptrace doctor --live
When a migration fails
up exits non-zero, and the error names the step file that failed. The report shows which steps completed and where the next run continues:
Apply 3 steps of 1 migration, group 2
# migration step result
1 20260922080949_demo_split_customer_name 1-expand.up.pg.sql completed 14ms
- restart required every process must run the new release
2 20260922080949_demo_split_customer_name 2-backfill.up.pg.go failed 2 batches, cursor=4
3 20260922080949_demo_split_customer_name 3-finalize.up.pg.sql not run
A rerun continues step 2 with cursor=4.
Take these steps in order:
- See where the run stopped:shell
uptrace migrate status --full - Fix the cause, for example disk space, a ClickHouse server that went down, or a lock (see below).
- Run
upagain. It skips every completed step:- a PostgreSQL expand step rolled back as a whole, so it runs again from the start;
- a PostgreSQL backfill or finalize file continues at the first statement that did not complete;
- a Go backfill continues at its last committed batch;
- a ClickHouse file sends every statement again.
Lock timeouts
A PostgreSQL schema change waits for table locks. To avoid holding back the queries that queue behind it, up waits at most --lock-timeout (2 seconds by default) for each lock, then retries up to --lock-retries times (10 by default) with a growing pause. Before each retry it logs the transactions that hold the lock. An idle in transaction session is a frequent cause.
When the retries run out, up fails with a lock timeout. Close the blocking session and run up again, or raise the limits:
uptrace migrate up --lock-timeout 10s --lock-retries 30
Mark a migration by hand
Use mark_applied only when you confirmed that the work of the migration is already in the database, for example after you applied it by hand. A marked step never runs:
uptrace migrate mark_applied --version 20260922080949
status cannot confirm the work for you: read the step files (uptrace migrate show --version ...) and check each object they create. When you are not sure, run up again instead.
To take a wrong mark back without running any SQL:
uptrace migrate mark_unapplied --version 20260922080949
Schema drift
Before it runs, up compares the live schemas with the schema that the newest applied migration expects. If someone changed the schema outside the migrations, up prints an error and runs nothing:
ERROR: the PostgreSQL schema differs from each schema that migration 20260926142252_slo_subject_unknown lists.
expected: sha256:6f81…
sha256:c9b8…
live: sha256:8345…
A change outside the migrations causes this.
No migration runs. To see each difference, run:
uptrace migrate diff
If you accept the differences, run the same command again with --allow-schema-drift:
uptrace migrate up --allow-schema-drift
The flag works with --plan and --version too.
See each difference:
uptrace migrate diff
diff builds scratch PostgreSQL and ClickHouse databases with the migrations your databases completed, prints each object that differs, and drops the scratch databases. The PostgreSQL user needs the CREATEDB privilege.
If you accept each difference, for example an index someone added by hand or a new major PostgreSQL version that prints a definition differently, continue with:
uptrace migrate up --allow-schema-drift
Roll back
Down steps restore the shape of the schema. They restore data only where a migration writes it back. If the data matters, restore the backups you took before the upgrade.
Preview the rollback first:
uptrace migrate down --plan
A bare down rolls back the newest work: a run that stopped, otherwise the last group of migrations that ran together. Pass --version to roll back one migration.
A rollback runs with the new binary, because only the new binary contains the down steps. Keep a copy of it before you install the old version. A rollback mirrors the upgrade: the finalize and backfill down steps run while the new version serves, the processes move back to the old version, and the expand down steps run last. The three cases work the same way.
Built-in restarts:
# 1. With the new binary installed, run the down steps that the new version needs.
uptrace migrate down --before-release-change
# 2. Install the old binary.
# 3. With the kept copy of the new binary, restart Uptrace on the old version and finish the rollback.
/path/to/new/uptrace migrate down --restart
Manual restarts:
# 1. With the new binary installed, run the down steps that the new version needs.
uptrace migrate down --before-release-change
# 2. Install the old binary, and restart every Uptrace process.
sudo systemctl restart uptrace
# 3. With the kept copy of the new binary, finish the rollback.
/path/to/new/uptrace migrate down
Kubernetes: run each down as a Job with the new image, and change the image back with the deploy tool between them:
1. Job uptrace migrate down --before-release-change
2. deploy tool change the image back to the old version, and wait until the pods are ready
3. Job uptrace migrate down
Do not roll back through helm rollback hooks. Helm runs the hooks of the old revision, and the old image does not contain the down steps of the new migrations.
Give both runs of down the same selection: the same --version, or none.
A rollback that did not finish blocks up. up refuses to start and names the down command that finishes it.
Command reference
| Command | What it does |
|---|---|
uptrace migrate status | Prints which migrations are applied, and where a step stopped. Writes nothing. |
uptrace migrate up | Runs every step of the unapplied migrations. |
uptrace migrate show | Prints the SQL that up (or down with --down) would send. Runs nothing. |
uptrace migrate down | Runs the down steps of one migration or of the newest work. |
uptrace migrate mark_applied | Records steps as completed without running them. |
uptrace migrate mark_unapplied | Deletes the step records of one migration without running any SQL. |
uptrace migrate diff | Compares the live schemas with freshly migrated scratch databases. |
uptrace migrate init | Creates the step record table. up does it too, so you rarely need it. |
uptrace migrate reset | Drops the databases and creates them again, empty. Run up after it. |
uptrace migrate truncate | Deletes every row, and keeps the tables and the migration records. |
Common flags:
| Flag | Commands | Meaning |
|---|---|---|
--plan | up, down | Print what the command would run, and run nothing. |
--version | up, down, show, mark_* | Select one migration by its 14-digit version. |
--restart | up, down | Restart the serve and worker processes at each restart point. Needs a supervisor such as systemd. |
--before-release-change | up, down | Stop before the point where the processes must switch versions, and exit zero. |
--wait | up, down | How long to wait for processes to restart or acknowledge a hold. |
--lock-timeout | up | How long one PostgreSQL statement waits for a lock before it retries (default 2s). |
--lock-retries | up | How many times a statement retries after a lock timeout (default 10). |
--allow-schema-drift | up | Run even when a live schema differs from what the migrations expect. |
--full | status | Print every migration and every step. |
--fail-on-unapplied | status | Exit non-zero when a migration is unapplied or a rollback did not finish. |
--database pg|ch | show, reset, truncate | Work on one database only. |
--down | show | Print the down steps instead of the up steps. |