The Strangler Fig Pattern for Incremental Modernization

Use the Strangler Fig pattern to replace legacy capabilities incrementally with routing, data ownership, observability, and rollback plans.

published: reading time: 11 min read author: GeekWorkBench
Quick Summary

The Strangler Fig pattern replaces a legacy system one capability at a time by routing selected traffic through a facade. This guide covers migration boundaries, data ownership, backfills, change streams, and rollback while old and new implementations coexist. It also explains rollout gates, production failure scenarios, and retirement criteria so teams can move traffic with evidence and keep the migration bridge temporary.

The Strangler Fig Pattern for Incremental Modernization

Introduction

A team does not need to replace an entire application at once. It can route catalog reads to a new implementation while cart and checkout still use the legacy system:

routes:
  catalog: replacement
  cart: legacy
  checkout: legacy

That routing facade gives the team a small migration slice to validate before moving writes or another capability. This guide covers when the Strangler Fig pattern fits, how to handle data ownership and rollback, and what to measure as traffic shifts.

When to Use / When Not to Use

Use it when a valuable legacy system has high rewrite risk and capabilities can be redirected independently. It helps teams learn from production while retaining a fallback.

It is less suitable when there is no routing boundary, capabilities share inseparable transactions, or supporting two implementations costs more than a planned replacement. If the source of truth is unclear or divergence cannot be reconciled, clarify ownership before cutover.

Core Concepts

Incremental routing

A facade or proxy directs requests to the old or new implementation. Route by capability first and expose decisions; avoid hidden routing branches in clients.

Define entry conditions, success measures, and exit conditions for each slice. Shadow reads, enable a small cohort, expand in steps, and remove the old route after the rollback window. Shadow mode must not repeat irreversible side effects.

Choose the seam around a business capability with a stable contract, such as catalog search or account preferences, rather than splitting by technical layer alone. Keep route selection in one place and make its precedence explicit: an operator override might take priority over a cohort rule, which takes priority over the default legacy route. Preserve request identity, authorization, and correlation IDs across the facade. If a request can be retried, keep the chosen implementation stable for that operation so one logical action does not bounce between systems.

Before expanding traffic, test both paths against the same contract and record which path served each request. Define what the facade does when the replacement times out: fail the request, or retry the legacy path only when the operation is safe to repeat. Silent fallback can hide a failing replacement and duplicate writes, so every fallback needs a metric and an explicit policy.

Data ownership and consistency

Routing does not migrate data. Choose an authority for every entity and operation. Dual writes can partially succeed, so prefer one writer with an outbox or change-data-capture feed. Track lag, define conflict handling, and reconcile business invariants.

For reads, backfill a snapshot, replay later changes, validate data, then redirect. For writes, state when ownership moves and how outstanding messages are handled. See change data capture for replication without application dual writes.

Two database writes in one request are not an atomic dual write just because they share a handler: a timeout can leave one committed and the other unknown. If a temporary dual-write phase is unavoidable, make each operation idempotent, persist enough information to retry or reconcile it, and alert on mismatches. Give the cutover a deadline. Once the new store owns writes, the old copy should be a derived replica rather than a second authority.

Rollback is a capability

Define rollback before routing traffic. Switching back is safe only if the old system can accept the new system’s data and side effects. This may require compatible schemas, reverse replication, or a forward-fix policy. A toggle alone is not a rollback plan.

During coexistence, keep contracts backward-compatible for the full period in which either implementation may receive traffic. Write down the rollback trigger, who can invoke it, and how to handle requests already in flight. If writes have crossed an irreversible boundary, such as sending a notification or capturing a payment, route-back may restore availability but cannot undo the effect; use idempotency keys and a compensating action where the domain allows one. After rollback, reconcile the interval before traffic moves again.

Mermaid Diagram

The router selects one owner for normal writes while a change stream supports migration and validation.

flowchart LR
    Client[Client] --> Router[Routing facade]
    Router -->|Legacy capability| Old[Legacy system]
    Router -->|Migrated capability| New[Replacement service]
    Old -->|Change stream| Replicator[Replication and validation]
    Replicator --> NewData[(New data store)]
    New --> NewData
    Old --> OldData[(Legacy store)]

Implementation or Decision Example

Suppose a commerce application handles catalog, cart, and checkout together. Start with catalog reads, which are high-volume and mostly read-only. Add a facade, backfill the new store, consume database changes, and compare counts, prices, and availability. Route staff traffic first, then a small customer cohort; watch stale reads and errors.

Keep checkout writes in the old application. Read success does not prove write safety. Record route decisions and service version by cohort. Expand after validation; if thresholds fail, route back, pause replication if needed, and reconcile changes made during rollout.

Trade-Off Table

Strategy Benefit Cost or risk
Big-bang replacement One final architecture and no long-lived bridge Concentrated release and data-conversion risk
Capability-by-capability routing Small releases and fast feedback Temporary operational and integration complexity
Shadow reads Compare outputs without changing user results Extra load; unsafe if reads trigger side effects
Dual writes Keeps both stores current in simple cases Partial success creates divergence
One writer plus change stream Clearer authority and replay path Replication lag and reconciliation work
Facade at the system edge Central route policy and consistent telemetry Can become a bottleneck or accumulate permanent special cases
Branch by abstraction inside the legacy code Useful when callers cannot be redirected at the network edge Requires safe internal seams and coordinated code changes

Production Failure Scenarios

Failure Effect Mitigation
Wrong routing rule Some calls hit an unready replacement Test routing policy, expose cohort metrics, and keep a kill switch
Replication lag rises Replacement serves stale data Alert on lag and offset, pause traffic increase, and define acceptable staleness
Dual write partially fails Stores disagree Use one authoritative writer, replayable outbox/change stream, and reconciliation
Rollback loses new writes Users see missing or old state Design reverse compatibility or stop writes and reconcile before switching
Bridge remains forever Both systems and routing logic need support Set capability retirement criteria, owners, and a deadline review
Replacement times out but fallback also runs A retry can duplicate a command or produce two side effects Limit fallback to idempotent operations, use a shared idempotency key, and make fallback visible in metrics
Identity or authorization mapping differs A user can lose access or receive access they should not have Compare authorization decisions before cutover and test migrated identities with representative roles
A consumer processes an event twice or out of order Derived state drifts even though replication appears healthy Make consumers idempotent, monitor sequence gaps, and provide replay plus invariant checks

Observability Checklist

  • Tag requests with implementation and rollout cohort.
  • Compare error rates, latency, and business outcomes between old and new paths.
  • Track backfill progress, replication offset, lag, and rejected events.
  • Reconcile domain invariants such as totals and uniqueness, not just row counts.
  • Alert on rollback switches and keep a timestamped audit of routing changes.
  • Document who can halt the rollout and how to recover unprocessed messages.

Security and Compliance Notes

Copies multiply sensitive data. Restrict snapshot, stream, log, and staging access; encrypt transfers and apply retention rules. Preserve deletion and consent behavior in both systems, and confirm residency and audit requirements after cutover.

Common Pitfalls / Anti-Patterns

  • Calling a full rewrite “incremental” when no capability can be cut over independently.
  • Routing some traffic to a replacement without a data authority model.
  • Using dual writes without idempotency, retry handling, or reconciliation.
  • Shadowing writes that send emails, charge cards, or create external effects twice.
  • Treating route-back as safe even after incompatible writes.
  • Leaving the router, replication job, and old system without retirement criteria.

Quick Recap Checklist

  • Is the chosen capability separable and observable at a stable boundary?
  • Is there one clear source of truth for each operation and entity?
  • Can data be backfilled, replayed, and reconciled?
  • Does rollback account for writes and side effects already made?
  • Are rollout cohorts, success thresholds, and retirement criteria explicit?

Interview Questions

1. Why is this pattern safer than a full rewrite?
It limits each release to a smaller capability and lets the team compare behavior in production before moving more traffic. It does not remove risk; it spreads risk across controlled steps with measurable gates.
2. How should data be handled while both systems exist?
Assign an authoritative writer for each operation, then use backfill and a replayable change stream when another store needs a copy. Define lag, conflict, and reconciliation rules before cutover.
3. What makes a rollback plan credible?
It accounts for data and side effects produced by the replacement. The old system must be able to consume them, or the team needs a reverse sync or an explicit forward-fix policy. A route toggle alone is not enough.
4. What is the role of a routing facade?
It gives operators one visible place to direct a capability to the legacy or replacement implementation. It should expose the route decision in telemetry and remain simple enough to test and retire.
5. How do you choose a capability boundary for migration?
Choose a business capability with a contract and data boundary that can be observed independently. Check its callers, transactions, and ownership dependencies; if moving it requires synchronized changes across many unrelated capabilities, first create a safer seam inside the existing system.
6. Why are dual writes risky, and what is a safer alternative?
The two writes can have different outcomes after a timeout or partial failure, leaving no reliable single truth. Prefer one authoritative writer and a transactional outbox or change-data-capture stream; if dual writes are temporary, use idempotency, retries, mismatch alerts, and reconciliation.
7. When is fallback from the replacement to the legacy system safe?
Fallback is safest for repeatable reads or idempotent commands when both systems share compatible state. It can duplicate effects or return stale data otherwise, so define eligible operations, preserve an idempotency key, and measure every fallback.
8. What should a rollout gate measure?
Use technical signals such as error rate, latency, and replication lag alongside business invariants such as totals, uniqueness, and successful task completion. Compare the replacement with a baseline over a representative cohort and set stop thresholds before expanding traffic.
9. How can a team prevent the migration facade from becoming permanent?
Assign an owner and retirement condition to each route, track remaining legacy traffic, and review the bridge as a time-bounded operational component. Remove a route only after consumers, data, and rollback windows have all moved or closed.
10. What is the difference between shadow traffic and a canary rollout?
Shadow traffic sends a copy to the replacement for comparison while the legacy response remains authoritative. A canary serves real user requests from the replacement for a limited cohort, so it tests user-visible behavior and needs a rollback threshold. Shadowing is unsafe for requests that cause side effects unless those effects are suppressed.

Further Reading

Conclusion

The Strangler Fig pattern works when a team can move one capability at a time and keep each transition observable. Define data authority before writes move, and make rollback account for changed state. The bridge is temporary only with clear ownership and retirement conditions.

Category

Related Posts

Architecture Decision Records: A Working Guide

Use concise architecture decision records to capture context, options, consequences, and revisit triggers so teams can understand design choices later.

#software-architecture #adr #technical-decisions

Architecture Styles and Patterns: A Practical Guide

Compare layered, hexagonal, event-driven, and service architectures using boundaries, deployment needs, failure modes, and a concrete selection method.

#software-architecture #architecture-patterns #system-design

Architecture Fitness Functions as Executable Guardrails

Turn architecture principles into executable fitness functions that check dependency rules, performance limits, and deployment constraints as systems evolve.

#software-architecture #architecture-testing #fitness-functions