Coordinating Database Migrations Across Services

Philip Rehberger Oct 9, 2026 7 min read

Sequence schema changes across multiple services without coupling deploys. Includes expand/contract orchestration.

When a single service owns its own database, schema migrations are a self-contained problem. The service runs the migration before (or after) deploying the code that needs it. Done.

The moment two services share a database — or one service consumes another service's API that just changed — migrations become a coordination problem. The wrong sequence creates downtime. The right sequence is rarely the most obvious one.

This post is the patterns for coordinating migrations across services without scheduled maintenance windows.

The Three Variants

Single service, single database. One team, one migration. Easy.

Multiple services, shared database. Multiple teams own different tables in the same database. Migrations affect everyone.

Multiple services, each with their own database, communicating via API. A schema change in one service's database might change its API contract.

The first case is what most tutorials cover. The other two are where production teams spend their time.

The Coupling Problem

A "schema migration" is rarely just a schema change. It is also:

  • The data already in the table
  • The code that reads and writes it
  • The API contracts that expose it
  • The consumers of that API

When a column is renamed, all of these have to agree on the new name. Coordinating the change is the actual work.

The Expand-Contract Pattern

The canonical pattern for backward-incompatible schema changes is expand-then-contract:

  1. Expand. Add the new structure alongside the old. Both work.
  2. Migrate. Move readers and writers to use the new structure.
  3. Contract. Remove the old structure once nothing uses it.

For a column rename:

-- Step 1: Expand — add new column
ALTER TABLE users ADD COLUMN full_name VARCHAR(255);

-- Step 2: Migrate — populate, then deploy code that writes both
UPDATE users SET full_name = first_name || ' ' || last_name;
-- (deploy code that writes to both columns)
-- (deploy code that reads from full_name)

-- Step 3: Contract — drop the old column
ALTER TABLE users DROP COLUMN first_name;
ALTER TABLE users DROP COLUMN last_name;

The expansion phase makes the migration safe. Both old and new code can run during the transition. Nothing breaks.

The contraction phase happens days or weeks later, once you are certain nothing reads the old columns.

Schema Changes Across Services

For a shared database with multiple services, expand-contract is essential. The orchestration:

  1. Coordinate the migration. A team writes the SQL; other teams agree on the timing.
  2. Apply the expand migration. Both old and new structures exist.
  3. Each service deploys its code change. No strict ordering required — both versions are valid.
  4. Verify all services are on the new code. Dashboards, deploy logs, metrics.
  5. Apply the contract migration. Old structure goes away.

The migration window can be days or weeks. The system is never broken during the window. Everyone deploys when they are ready.

API Versioning Plus Schema Migrations

For services that own their own database and expose an API, schema changes often imply API changes. The coordination shifts:

  1. The service's API gains a new version that exposes the new shape (/v2/users/me returning full_name).
  2. Consumers migrate to v2 at their own pace.
  3. Once all consumers are on v2, v1 can be deprecated.
  4. Once v1 is removed, the underlying schema can contract.

Each consumer's migration to v2 is independent. The provider does not block on any single consumer.

The Data Backfill Problem

Schema changes are easy. Data backfills are where teams trip up.

Adding full_name to a 100-million-row table:

  • UPDATE users SET full_name = ... locks every row, takes hours, kills the database.
  • The same query in chunks of 10,000 rows with sleeps between, runs over a weekend.

For large tables, the pattern is:

  1. Add the column (instant, no rewrite in modern databases).
  2. Backfill in batches asynchronously.
  3. Verify the backfill is complete before deploying code that depends on the new column being populated.

Most backfills can be done as a background job:

class BackfillFullName implements ShouldQueue
{
    public function handle(): void
    {
        User::whereNull('full_name')
            ->limit(10000)
            ->each(function ($u) {
                $u->update(['full_name' => trim("{$u->first_name} {$u->last_name}")]);
            });

        if (User::whereNull('full_name')->exists()) {
            self::dispatch()->delay(now()->addSeconds(10));
        }
    }
}

The job processes 10,000 rows per run, with a brief pause, until done. Coexists with normal write traffic.

Coordinating Across Deploys

When multiple deploys must happen in sequence — schema, then service A, then service B — the coordination becomes the migration's hardest part.

Three patterns:

Manual sequencing. A release manager (or runbook) coordinates the steps. Slow, error-prone, but works for infrequent migrations.

Feature flags. New code paths are flag-gated. Deploy code first with flag off; deploy schema; turn flag on. The flag becomes the coordination point.

Migration framework with dependency declarations. Tools like Liquibase, Flyway, or Atlas track migration state. A service's code waits to start until its required migrations have run.

For most teams, feature flags are the lightest-weight option. The deploys themselves are independent; the flag controls when the new path activates.

Long-Running Migrations

Some migrations take hours or days. A 500-million-row table being re-indexed, a foreign key being added, a column type being changed.

Tools that make this practical:

  • pt-online-schema-change (Percona) — performs MySQL schema changes without locking. Works by creating a shadow table, copying data, swapping at the end.
  • gh-ost (GitHub) — similar to pt-osc but uses the binary log instead of triggers.
  • pg_repack — PostgreSQL equivalent for table rewrites.
  • Vitess Online DDL — for sharded MySQL.

These tools turn a 4-hour table lock into a 4-hour background process that runs without blocking traffic.

For PostgreSQL specifically, modern versions support many operations online that older versions required downtime for. Adding columns with defaults, adding indexes concurrently, and many type changes are now non-blocking. Check the database's documentation before assuming a migration needs special tooling.

Rolling Back

Migrations can fail. The rollback story matters.

Three classes:

  • Trivially reversible. DROP TABLE reverses CREATE TABLE. Easy.
  • Reversible with data risk. Dropping a column means data loss. The reverse migration restores the column but the data is gone.
  • Effectively irreversible. A migration that produces a new derived state from the old data. Reversing requires reconstructing the original.

For the third class, rollback through forward migrations is the safer pattern. Instead of "undo migration 47," create migration 48 that returns the system to a workable state.

Always test the rollback in staging. The number of migrations that fail in rollback because nobody tested it is significant.

Multi-Region Migrations

For systems with multi-region databases (replicated or independently regional), migrations have to handle the multi-region case:

  • Migration must apply to all regions
  • Code change must be compatible with all regions during the rolling update
  • Replication delays must be considered

For multi-master systems, schema changes typically apply to one region first and propagate. Until all regions have applied, both schemas coexist.

For active-passive multi-region, the primary's schema migration replicates to standbys. Code deploys to all regions only after all replicas are caught up.

Common Mistakes

  • Big-bang migrations. Schema + code + data all change at once. One thing fails; everything rolls back.
  • No backfill plan. Adding a NOT NULL column to a 100M-row table with no plan for the existing rows.
  • Skipping the expand-contract pattern. Renaming a column directly. Production breaks during the deploy window.
  • No verification step. Assuming a migration ran successfully without checking.
  • No staging rehearsal. First time the migration runs is in production.

A Practical Process

For a non-trivial migration:

  1. Write the migration in expand-contract phases
  2. Test in staging with production-shaped data
  3. Plan the deploy sequence (which services, in which order)
  4. Identify the rollback path for each phase
  5. Run the expand phase
  6. Deploy services
  7. Verify (dashboards, sample queries, audit logs)
  8. Run the contract phase later — after a buffer period

A complex migration takes weeks from design to contract phase. The buffer between expand and contract is intentional; rushing it produces incidents.


Working on a schema change that touches multiple services or has to coordinate across teams? We help teams plan and execute migrations without scheduled downtime. scopeforged.com

Share this article

Related Articles

Need help with your project?

Let's discuss how we can help you build reliable software.