Upgrading Uptrace

Every Uptrace release can change the PostgreSQL and ClickHouse schemas. The uptrace migrate command applies those changes and moves the running processes to the new version, so Uptrace keeps accepting telemetry during the upgrade.

The short version for a DEB or RPM install:

shell
# 1. Back up PostgreSQL and ClickHouse.
# 2. Install the new package (do not restart the service yet).
# 3. Rewrite deprecated config options, then check the config.
uptrace config fix
uptrace config fix --write
uptrace config validate
# 4. Preview the upgrade.
uptrace migrate up --plan
# 5. Apply the migrations and restart Uptrace.
uptrace migrate up --restart

The rest of this page explains each step, what to do when a step fails, and how to roll back.

Supported upgrade paths

You can upgrade from each of these releases directly to the latest release, in one step. You do not need to install the releases between them:

You runUpgrade
v2.1.0 or laterFollow this page.
v2.0.0-beta.3 to v2.0.3, v2.1.0-beta to v2.1.0-beta.8Read Upgrading from v2.0 or a v2.1 beta first, then follow this page.
v2.0.0-beta.1, v2.0.0-beta.2, v2.0.0-rc.2Not supported. Install the latest release on new databases.
v1.xNot supported. v2 stores data differently. Run v2 next to v1 as the v2.0 migration strategy describes, then remove v1.

Every release in the table is tested: its own binary creates the databases, and the latest release then upgrades them.

Before you upgrade

Before you start:

  • Read the release notes on GitHub Releases.
  • Back up both PostgreSQL and ClickHouse. The Docker, Kubernetes, and Ansible guides show how.
  • With the new binary installed, check that it can start against your stores:
    shell
    uptrace preflight
    

    preflight checks that the binary has the license keys of its edition, that PostgreSQL, every ClickHouse replica, and the Kafka brokers (when configured) answer with your credentials, that PostgreSQL and ClickHouse run supported versions, and that no migration stopped partway and no rollback was left unfinished. It exits non-zero when one of these fails. Unapplied migrations are expected at this point, so they only warn. preflight runs a subset of uptrace doctor, which also checks Redis, the replicas, the disks, and the outbound endpoints.

All uptrace migrate commands read /etc/uptrace/config.yml by default. Pass --config=/path/to/config.yml to use another file.

Upgrading from v2.0 or a v2.1 beta

v2.0.x and the v2.1.0 betas are classic releases: they migrated the databases with other commands and recorded the migrations in other tables. The latest release upgrades their databases, but the first upgrade needs these extra steps.

  1. Upgrade the databases. Uptrace now needs ClickHouse 26.3 or later and PostgreSQL 15 or later. Upgrade the servers before you install the new Uptrace: v2.0.x and the v2.1 betas keep working on ClickHouse 26.3. uptrace preflight checks both versions.
  2. Set a secret. The new release refuses to start when service.secret is the placeholder FIXME that older configs shipped with:
    text
    invalid config: service.secret is the shipped placeholder "FIXME"; set a random value
    

    Set it to a random value:
    shell
    openssl rand -hex 32
    

    The secret signs login sessions, so a new secret logs out every user once. The DEB and RPM packages replace the placeholder for you.
  3. Rewrite the config. Run uptrace config fix --write. It moves the options that changed place, for example spans to pipelines.spans and ch_schema.metrics to ch_schema.timeseries and ch_schema.datapoints, and deletes the options that the new release does not use.
  4. Replace the old migration commands in your scripts, init containers, and entrypoints. serve no longer migrates the databases when it starts, as v2.0 did:
    Old commandNew command
    uptrace pg init, uptrace pg migrateuptrace migrate up
    uptrace ch init, uptrace ch migrateuptrace migrate up
    uptrace ch status, uptrace ch checkuptrace migrate status
    uptrace sync_dashboardsuptrace dashboard sync

    uptrace migrate up migrates PostgreSQL and ClickHouse together. The old commands print the new command and exit with an error.
  5. Restart the old processes yourself. A classic process cannot receive the restart request of --restart. After the expand steps, uptrace migrate up checks PostgreSQL for the sessions of a classic process. While one is connected, up waits up to --wait, then stops before the backfill, names the sessions, and exits non-zero. The old release keeps serving. Use the manual restarts case for the first upgrade:
    shell
    uptrace migrate up --before-release-change
    sudo systemctl restart uptrace
    uptrace migrate up
    

    Later upgrades can use --restart.

On the first run, migrate reads the records of the classic release and does not run those migrations again. Before that run, status and preflight count every migration as unapplied. up --plan tells you how many steps it will take over:

text
The classic runners applied 7 steps that have no step record. They read as applied, and the next up, mark_applied, mark_unapplied, or down records them.

up then prints Adopted 7 steps that the classic runners applied. and runs the rest. It leaves the old bun_migrations and ch_migrations tables in place.

The DEB and RPM packages do not restart Uptrace onto databases that still need migrations. The old process keeps serving until you run the commands above.

How migrations run

A migration can have up to three phases, and the processes switch to the new version in the middle:

text
old version serves traffic
   │  expand     add new tables, columns, and indexes; the old version keeps working
   ▼
restart: every Uptrace process moves to the new version
hold: some migrations pause a capability, for example alerts, while the next phases run
   │  backfill   copy and convert existing rows into the new shape
   │  finalize   remove what only the old version needed; build the final indexes
   ▼
release: the paused capabilities start again
  • Expand runs while the old version still serves, so it only adds objects the old version can ignore.
  • Restart required is the point where every serve and worker process must run the new binary before the work continues.
  • Hold and release pause and resume a capability for the duration of the backfill and finalize phases. No process restarts for them.
  • Backfill and finalize run on the new version. A long backfill commits in batches, so a stop does not lose its progress.

uptrace migrate records every completed step in PostgreSQL. A second run skips completed steps and continues where the previous run stopped.

Preview the upgrade

Use three read-only commands before you change anything. None of them takes a lock or writes to the databases.

See what the databases hold now:

shell
uptrace migrate status
text
Status     behind
Applied    41
Unapplied  2

  20260915101500_heatmap     unapplied
  20260916090000_retention   unapplied

Run: uptrace migrate up --plan

See what up will do, in order:

shell
uptrace migrate up --plan
text
Plan  2 steps of 1 migration

Migrations
  20260922080949_demo_split_customer_name  split the customer name into first and last

  #  migration                                step                  runs as
  -  restart required                                               every process must run the new release
  -  hold                                                           every process stops alerts and acks
  1  20260922080949_demo_split_customer_name  2-backfill.up.pg.go   1 transaction per batch, resumes with cursor=4
  2  20260922080949_demo_split_customer_name  3-finalize.up.pg.sql  6 statements, no transaction; a stop resumes the file
  -  release                                                        every process starts alerts again

Completed, not run again
     migration                                step                completed at
     20260922080949_demo_split_customer_name  1-expand.up.pg.sql  2026-09-22 08:14:03 UTC, group 2

Transactions      1 per batch of step 1
Restart required  once: before step 1
Holds             alerts: held before step 1, released after step 2

In this example, an earlier run already completed the expand step, so the plan continues the backfill at its last batch. The Restart required line tells you whether the processes must switch versions, and the Holds line tells you which capabilities pause and for how long. Add --plan to the exact command you intend to run: the same flags select the same work.

To read the SQL itself, print it without running it:

shell
uptrace migrate show
uptrace migrate show --database pg
uptrace migrate show --database ch

Choose how the processes restart

Every upgrade runs in the same order:

  1. The expand steps run while the old version serves.
  2. Every serve and worker process moves to the new version.
  3. The backfill and finalize steps run on the new version.

What differs is who moves the processes in step 2:

CaseWho restarts the processesCommandsUse it with
Built-in restartsuptrace migrateup --restartsystemd, Ansible
Manual restartsYouup --before-release-change, then restart, then upsystemd, any supervisor
KubernetesYour deploy toolThe same as manual restarts, run as JobsHelm, Argo CD

Manual restarts and Kubernetes are the same case: up stops before step 2, something else changes the version, and a second up finishes the work. On Kubernetes, the deploy tool changes the image, and the two up runs are Jobs.

Use the same steps for a release that adds no migration. Then up runs nothing, and only the processes restart.

Case 1: built-in restarts

uptrace migrate up --restart does the whole upgrade in one run. It needs a supervisor that starts a stopped process again from the installed binary. The packaged systemd unit sets Restart=always, so this works for the DEB, RPM, and binary installs. The Ansible playbook uses this case.

  1. Install the new version using DEB, RPM, or a pre-compiled binary. Do not restart the service.
  2. Fix and validate the config. A new release can move or rename config options. The old keys keep working, but Uptrace logs a deprecation warning for each one. Review the fixes, apply them, and then validate the result:
    shell
    uptrace config fix
    uptrace config fix --write
    uptrace config validate
    

    config fix without --write prints each fix and a diff, and changes nothing. --write rewrites the file and keeps the original as config.yml.bak. See Fixing deprecated options.
  3. Apply the migrations and restart Uptrace:
    shell
    uptrace migrate up --restart
    

up --restart runs the expand steps while the old version serves, then asks each running process to stop. systemd starts it again from the new binary. Workers restart first, then each serve process one at a time, so the API keeps serving when you run more than one. After that, up runs the backfill and finalize steps. At the end, it restarts every process that the run did not restart yet, so each process reads the new binary, config, and data files.

Use the same command after you change only the config or a data file: it restarts every process even when no migration is left.

--wait bounds how long up waits for each restart (2 minutes by default):

shell
uptrace migrate up --restart --wait 5m

If a process does not come back in time, up stops and names it. Fix the cause, for example a host where the new binary is not installed, and run the same command again.

Case 2: manual restarts

Use this case when you restart the processes yourself, for example with systemctl on each host, or with a supervisor that --restart cannot use.

  1. Install the new version without a restart. Run uptrace config fix --write to rewrite deprecated options, and run uptrace config validate.
  2. Run the expand steps. up stops before the restart and exits zero. The old version keeps serving:
    shell
    uptrace migrate up --before-release-change
    
  3. Restart every Uptrace process on every host, so it runs the new binary:
    shell
    sudo systemctl restart uptrace
    
  4. Run the backfill and finalize steps:
    shell
    uptrace migrate up
    

Before the backfill, step 4 checks that every live process runs the new version. If one still runs the old version, up stops and names it. Restart that process and run up again, or pass --wait 10m to let up wait for the restarts:

shell
uptrace migrate up --wait 10m

Case 3: Kubernetes

Kubernetes is case 2, with the deploy tool in step 3. The kubelet starts a stopped container again from the image of its pod spec, so only the deploy tool can move a pod to the new version. Run each up as a Job with the new image:

text
1. pre-upgrade Job    uptrace migrate up --before-release-change
2. deploy tool        change the image, and wait until the new pods are ready
3. post-upgrade Job   uptrace migrate up --wait 10m
  • Do not use --restart on Kubernetes. It would only restart each pod again on the same image.
  • Start step 3 only after the rollout completes. With Helm, run helm upgrade --wait, so Helm runs the post-upgrade hooks after the resources are ready. With Argo CD, use a PostSync hook. If a pod still runs the old image, step 3 waits up to --wait for it, then stops and names the pod. Run the Job again after the rollout.

The Helm chart runs these steps for you. A pre-upgrade Job runs step 1 with the new image and the new config, helm upgrade changes the image in step 2, and a post-upgrade Job runs step 3. If a Job fails, the upgrade stops, and the old pods keep serving. See upgrading with Helm.

Docker

Pull the new images and recreate the containers as described in upgrading with Docker.

Check the result

When the upgrade completes, status reports that the databases are up to date:

shell
uptrace migrate status
text
Status   up to date
Applied  43

In a script or a health check, use --fail-on-unapplied. It exits non-zero when a migration is unapplied or a rollback did not finish:

shell
uptrace migrate status --fail-on-unapplied

To check what the upgraded processes see, read the report of each running process. It fails when a process serves with migrations that are not applied:

shell
uptrace doctor --live

When a migration fails

up exits non-zero, and the error names the step file that failed. The report shows which steps completed and where the next run continues:

text
Apply  3 steps of 1 migration, group 2

  #  migration                                step                  result
  1  20260922080949_demo_split_customer_name  1-expand.up.pg.sql    completed  14ms
  -  restart required                                               every process must run the new release
  2  20260922080949_demo_split_customer_name  2-backfill.up.pg.go   failed     2 batches, cursor=4
  3  20260922080949_demo_split_customer_name  3-finalize.up.pg.sql  not run

A rerun continues step 2 with cursor=4.

Take these steps in order:

  1. See where the run stopped:
    shell
    uptrace migrate status --full
    
  2. Fix the cause, for example disk space, a ClickHouse server that went down, or a lock (see below).
  3. Run up again. It skips every completed step:
    • a PostgreSQL expand step rolled back as a whole, so it runs again from the start;
    • a PostgreSQL backfill or finalize file continues at the first statement that did not complete;
    • a Go backfill continues at its last committed batch;
    • a ClickHouse file sends every statement again.

Lock timeouts

A PostgreSQL schema change waits for table locks. To avoid holding back the queries that queue behind it, up waits at most --lock-timeout (2 seconds by default) for each lock, then retries up to --lock-retries times (10 by default) with a growing pause. Before each retry it logs the transactions that hold the lock. An idle in transaction session is a frequent cause.

When the retries run out, up fails with a lock timeout. Close the blocking session and run up again, or raise the limits:

shell
uptrace migrate up --lock-timeout 10s --lock-retries 30

Mark a migration by hand

Use mark_applied only when you confirmed that the work of the migration is already in the database, for example after you applied it by hand. A marked step never runs:

shell
uptrace migrate mark_applied --version 20260922080949

status cannot confirm the work for you: read the step files (uptrace migrate show --version ...) and check each object they create. When you are not sure, run up again instead.

To take a wrong mark back without running any SQL:

shell
uptrace migrate mark_unapplied --version 20260922080949

Schema drift

Before it runs, up compares the live schemas with the schema that the newest applied migration expects. If someone changed the schema outside the migrations, up prints an error and runs nothing:

text
ERROR: the PostgreSQL schema differs from each schema that migration 20260926142252_slo_subject_unknown lists.
  expected: sha256:6f81…
            sha256:c9b8…
  live:     sha256:8345…
A change outside the migrations causes this.
No migration runs. To see each difference, run:
  uptrace migrate diff
If you accept the differences, run the same command again with --allow-schema-drift:
  uptrace migrate up --allow-schema-drift
The flag works with --plan and --version too.

See each difference:

shell
uptrace migrate diff

diff builds scratch PostgreSQL and ClickHouse databases with the migrations your databases completed, prints each object that differs, and drops the scratch databases. The PostgreSQL user needs the CREATEDB privilege.

If you accept each difference, for example an index someone added by hand or a new major PostgreSQL version that prints a definition differently, continue with:

shell
uptrace migrate up --allow-schema-drift

Roll back

Down steps restore the shape of the schema. They restore data only where a migration writes it back. If the data matters, restore the backups you took before the upgrade.

Preview the rollback first:

shell
uptrace migrate down --plan

A bare down rolls back the newest work: a run that stopped, otherwise the last group of migrations that ran together. Pass --version to roll back one migration.

A rollback runs with the new binary, because only the new binary contains the down steps. Keep a copy of it before you install the old version. A rollback mirrors the upgrade: the finalize and backfill down steps run while the new version serves, the processes move back to the old version, and the expand down steps run last. The three cases work the same way.

Built-in restarts:

shell
# 1. With the new binary installed, run the down steps that the new version needs.
uptrace migrate down --before-release-change

# 2. Install the old binary.

# 3. With the kept copy of the new binary, restart Uptrace on the old version and finish the rollback.
/path/to/new/uptrace migrate down --restart

Manual restarts:

shell
# 1. With the new binary installed, run the down steps that the new version needs.
uptrace migrate down --before-release-change

# 2. Install the old binary, and restart every Uptrace process.
sudo systemctl restart uptrace

# 3. With the kept copy of the new binary, finish the rollback.
/path/to/new/uptrace migrate down

Kubernetes: run each down as a Job with the new image, and change the image back with the deploy tool between them:

text
1. Job           uptrace migrate down --before-release-change
2. deploy tool   change the image back to the old version, and wait until the pods are ready
3. Job           uptrace migrate down

Do not roll back through helm rollback hooks. Helm runs the hooks of the old revision, and the old image does not contain the down steps of the new migrations.

Give both runs of down the same selection: the same --version, or none.

A rollback that did not finish blocks up. up refuses to start and names the down command that finishes it.

Command reference

CommandWhat it does
uptrace migrate statusPrints which migrations are applied, and where a step stopped. Writes nothing.
uptrace migrate upRuns every step of the unapplied migrations.
uptrace migrate showPrints the SQL that up (or down with --down) would send. Runs nothing.
uptrace migrate downRuns the down steps of one migration or of the newest work.
uptrace migrate mark_appliedRecords steps as completed without running them.
uptrace migrate mark_unappliedDeletes the step records of one migration without running any SQL.
uptrace migrate diffCompares the live schemas with freshly migrated scratch databases.
uptrace migrate initCreates the step record table. up does it too, so you rarely need it.
uptrace migrate resetDrops the databases and creates them again, empty. Run up after it.
uptrace migrate truncateDeletes every row, and keeps the tables and the migration records.

Common flags:

FlagCommandsMeaning
--planup, downPrint what the command would run, and run nothing.
--versionup, down, show, mark_*Select one migration by its 14-digit version.
--restartup, downRestart the serve and worker processes at each restart point. Needs a supervisor such as systemd.
--before-release-changeup, downStop before the point where the processes must switch versions, and exit zero.
--waitup, downHow long to wait for processes to restart or acknowledge a hold.
--lock-timeoutupHow long one PostgreSQL statement waits for a lock before it retries (default 2s).
--lock-retriesupHow many times a statement retries after a lock timeout (default 10).
--allow-schema-driftupRun even when a live schema differs from what the migrations expect.
--fullstatusPrint every migration and every step.
--fail-on-unappliedstatusExit non-zero when a migration is unapplied or a rollback did not finish.
--database pg|chshow, reset, truncateWork on one database only.
--downshowPrint the down steps instead of the up steps.