Post

Cloud-Native and Microservices: Scale, Resilience, and the 12-Factor App

Cloud-Native and Microservices: Scale, Resilience, and the 12-Factor App

TL;DR

  • Cloud-native architecture builds systems that run reliably on distributed infrastructure by following the 12-factor methodology — regardless of language or framework.
  • Microservices are independently deployable units — use them only when you’ve outgrown a monolith.
  • Resiliency patterns (circuit breakers, retries, idempotency) are essential for services that call other services over a network — and they’re exactly the kind of subtle, stateful logic that AI agents get plausible-looking but subtly wrong.

Why This Post Exists in the AI-Assisted Development Series

The previous post was about structure inside a single running process: which layer talks to which. This post is about what happens once that process has to survive on a network, next to other processes, some of which will fail.

That distinction matters more, not less, when an agent is doing a lot of the typing. A single-process layering violation is usually visible in a diff. A twelve-factor violation — a hardcoded connection string, a piece of state kept in memory instead of in a shared store, a missing timeout on a network call — often isn’t. It looks completely normal in isolation and only breaks when you run three copies of the process behind a load balancer, or redeploy at 2 a.m. and lose in-memory session data. These are exactly the failure modes an agent, working file-by-file without a live picture of your production topology, cannot be expected to catch on its own.

So this post gives you two things to bring into any AI-assisted work on a distributed system:

  1. A checklist (the 12 factors) you can hand to an agent as an explicit spec, instead of hoping “make it cloud-native” gets interpreted correctly.
  2. A small set of resilience patterns (circuit breaker, retry-with-backoff) you should recognize on sight, because agents readily produce a version of them that looks right, runs once successfully, and fails the first time it meets a real, sustained outage.

Prerequisites

  • Architecture Intro
  • Basic understanding of HTTP and REST APIs
  • Comfort with exception/error handling and basic concurrency concepts in any language

Concept Explanation

Cloud-native means building applications that take full advantage of cloud infrastructure — horizontal scaling, managed services, containerisation, and continuous delivery. A cloud-native app doesn’t just run in the cloud; it’s designed for the cloud’s characteristics: distributed failure, elastic scaling, and network-based service communication.

Microservices is one approach to cloud-native architecture: breaking a system into small, independently deployable services that each own a specific domain. Contrast with a monolith (one deployable unit), which is simpler to build and operate until scale demands more.

Rule of thumb: Start with a well-structured monolith. Migrate to microservices when you have specific, observed bottlenecks — not because microservices sound more professional, and not because an agent suggested splitting things up.


How It Works: When to Split

graph LR
    M[Well-Structured Monolith]
    M -->|Team grows, deploy conflicts| Split1[Extract Auth Service]
    M -->|Data volume, query load| Split2[Extract Analytics Service]
    M -->|Regulatory isolation| Split3[Extract Payment Service]
    M -->|Everything else| Keep[Stay in Monolith]

The 12-Factor App Principles

The 12-factor methodology defines how cloud-native applications should be built. Each factor solves a specific class of production problem, and — importantly for this series — each factor is a concrete, checkable statement you can put in front of an agent or a reviewer.

#FactorWhat it meansCommon violation
1CodebaseOne repo, many deploysMultiple repos for one app
2DependenciesDeclare and isolate dependencies explicitly (a manifest file, a lockfile)Relying on whatever happens to be installed on the machine
3ConfigStore in environment variables, never in codeHardcoded API keys, DB URLs in source
4Backing servicesTreat DB, cache, queue as attached resourcesAssuming a local database in prod code
5Build, Release, RunSeparate build and run stagesBuilding artefacts at runtime
6ProcessesStateless, share-nothing processesIn-memory session state
7Port bindingExport services via a portRequiring a separate web server host to be present
8ConcurrencyScale via process modelThreads/workers that share mutable global state
9DisposabilityFast startup, graceful shutdown60-second boot time, no shutdown-signal handler
10Dev/Prod ParityKeep environments as similar as possibleA lightweight file-based database in dev, a full server in prod
11LogsTreat as event streams, not filesWriting logs directly to a local file on disk
12Admin processesRun one-off tasks as process instancesScheduled cron logic baked inside the main app process

Most violated factors in AI-generated code: #3 (hardcoded config — an agent will happily inline a placeholder connection string “for now”), #6 (in-process state — the fastest correct-looking way to remember something between two requests is a global variable, which is also the wrong way), and #10 (a lightweight embedded database standing in for the real one in dev, with no equivalence check against prod).


Pseudocode Examples

1. Factor #3 — Twelve-Factor Config

record AppConfig:
    db_url: string
    cache_url: string
    api_key: string
    debug: boolean
    port: integer

function AppConfig.from_env():
    # Load all config from environment. Fail loudly if required vars are missing.
    required = ["DATABASE_URL", "CACHE_URL", "EXTERNAL_API_KEY"]
    missing = [key for key in required if env.get(key) is empty]
    if missing is not empty:
        raise EnvironmentError(
            "Missing required environment variables: " + join(missing, ", ") +
            "\nCopy .env.example to .env and fill in the values."
        )
    return AppConfig(
        db_url = env["DATABASE_URL"],
        cache_url = env["CACHE_URL"],
        api_key = env["EXTERNAL_API_KEY"],
        debug = lowercase(env.get("DEBUG", "false")) == "true",
        port = to_integer(env.get("PORT", "8000")),
    )

# Usage — config is loaded once at startup, injected everywhere
# config = AppConfig.from_env()

2. Microservice — Stock Quote Service

A minimal, independently deployable service following 12-factor principles. This is written as pseudocode against a generic “web framework” so you can map it onto Express, FastAPI, Spring Boot, ASP.NET Core, or anything else with the same shape: a router, a handler, a health check.

# stock_quote_service — Single-responsibility microservice.
# Exposes: GET /quote/{symbol}

configure_logging(format = "structured-json", destination = stdout)  # Factor 11
logger = get_logger()

app = new WebApp()

record StockQuote:
    symbol: string
    price: number
    timestamp: number

function fetch_price(symbol):
    # In production: call the real market-data provider using injected config.
    prices = {"RELIANCE": 2450.75, "TCS": 3600.00, "INFY": 1450.50}
    if symbol not in prices:
        raise NotFoundError("Symbol not found: " + symbol)
    return prices[symbol]

app.route("GET", "/quote/:symbol", function(request):
    symbol = uppercase(trim(request.params.symbol))
    try:
        price = fetch_price(symbol)
        quote = StockQuote(symbol, price, current_time())
        logger.info("Quote served: {} = {:.2f}", symbol, price)
        return json_response(quote, status = 200)
    catch NotFoundError:
        logger.warn("Unknown symbol requested: {}", symbol)
        return json_response({error: "Symbol not found"}, status = 404)
)

app.route("GET", "/health", function(request):
    # Liveness check — required for Kubernetes/ECS-style health probes.
    return json_response({status: "ok"}, status = 200)
)

port = to_integer(env.get("PORT", "5001"))  # Factor 7: port binding
app.listen(host = "0.0.0.0", port = port)   # 0.0.0.0 for container networking

3. Circuit Breaker — Resilience Pattern

When Service A calls Service B over the network, Service B can fail. Without a circuit breaker, Service A queues up requests, exhausting threads and memory until it also fails. The circuit breaker detects failures and short-circuits calls before they happen.

enum CircuitState:
    CLOSED     # Healthy: requests flow through
    OPEN       # Failing: requests blocked immediately
    HALF_OPEN  # Probing: one trial request allowed

class CircuitBreakerOpenError extends Error:
    # Raised when calls are attempted while circuit is OPEN.

class CircuitBreaker:
    # Transitions: CLOSED → OPEN after failure_threshold consecutive failures.
    # OPEN → HALF_OPEN after recovery_timeout has elapsed.
    # HALF_OPEN → CLOSED on success, → OPEN on failure.

    constructor(failure_threshold = 5, recovery_timeout = 30.0, name = "default"):
        self.name = name
        self.failure_threshold = failure_threshold
        self.recovery_timeout = recovery_timeout
        self.state_value = CLOSED
        self.failure_count = 0
        self.last_failure_time = null
        self.lock = new Mutex()

    function state():
        # Returns current state, transitioning OPEN→HALF_OPEN if timeout elapsed.
        with self.lock:
            if self.state_value == OPEN
               and self.last_failure_time is not null
               and (monotonic_time() - self.last_failure_time) > self.recovery_timeout:
                self.state_value = HALF_OPEN
            return self.state_value

    function call(func, ...args):
        current_state = self.state()

        if current_state == OPEN:
            raise CircuitBreakerOpenError(
                "Circuit '" + self.name + "' is OPEN — service unavailable. " +
                "Retry after " + self.recovery_timeout + "s."
            )

        try:
            result = func(...args)
            self._on_success()
            return result
        catch Exception as exc:
            self._on_failure()
            raise exc

    function _on_success():
        with self.lock:
            self.failure_count = 0
            self.state_value = CLOSED

    function _on_failure():
        with self.lock:
            self.failure_count += 1
            self.last_failure_time = monotonic_time()
            if self.failure_count >= self.failure_threshold:
                self.state_value = OPEN


# ── Usage ──────────────────────────────────────────────────────────────
function call_external_api(symbol):
    # Simulate a flaky external service.
    raise ConnectionError("Service temporarily unavailable")

breaker = new CircuitBreaker(failure_threshold = 3, recovery_timeout = 30, name = "Market-Data-API")

for attempt in range(1, 6):
    try:
        result = breaker.call(call_external_api, "RELIANCE")
    catch CircuitBreakerOpenError as e:
        print("Attempt " + attempt + ": Circuit OPEN — " + e.message)
    catch ConnectionError as e:
        print("Attempt " + attempt + ": Call FAILED — " + e.message + " | State: " + breaker.state())

4. Retry with Exponential Backoff

function retry_with_backoff(func, max_attempts = 3, base_delay = 1.0, max_delay = 30.0, jitter = true):
    # Exponential backoff: delay doubles each retry.
    # Jitter: adds randomness to prevent a thundering herd.
    for attempt in range(1, max_attempts + 1):
        try:
            return func()
        catch Exception as exc:
            if attempt == max_attempts:
                raise exc  # re-raise on final attempt

            delay = min(base_delay * (2 ^ (attempt - 1)), max_delay)
            if jitter:
                delay = delay * (0.5 + random() * 0.5)  # 50–100% of delay

            print("Attempt " + attempt + " failed: " + exc.message + ". Retrying in " + delay + "s...")
            sleep(delay)

Caching Architecture

graph LR
    Client -->|Request| API[API Service]
    API -->|Cache HIT: return instantly| Cache[(Distributed Cache)]
    API -->|Cache MISS: fetch + store| DB[(Database / External API)]
    DB -->|Data| API
    API -->|Response| Client

Cache levels by latency:

LevelTechnology examplesLatencyTTL Strategy
In-processAn in-memory map / built-in memoization~0.1µsProcess lifetime
DistributedRedis, Memcached~1msTTL per key type
CDNCloudFront, Cloudflare~10msCache-Control headers
Database queryNative query cache~10msInvalidated on write

AI in Development

Worked Example 1 — Generating 12-Factor-Compliant Code

The prompt:

1
2
3
4
5
6
7
Generate a microservice that follows 12-factor principles:
- Factor 3: All config (DB URL, API key, port) must come from environment
  variables, not hardcoded
- Factor 6: No in-memory state between requests (no global mutable variables)
- Factor 7: Bind to a port from a PORT environment variable
- Factor 11: Log as structured JSON to stdout, never to files
The service should expose: GET /health (liveness check) and GET /quote/{symbol}

What the agent produced (excerpt, pseudocode reflecting the actual shape):

API_KEY = "sk-live-4f9a2c..."      # <- hardcoded, not read from env
CACHE = {}                          # <- module-level mutable dict, shared across requests

app.route("GET", "/quote/:symbol", function(request):
    symbol = request.params.symbol
    if symbol in CACHE:
        return json_response(CACHE[symbol])
    price = fetch_price(symbol, api_key = API_KEY)
    CACHE[symbol] = price            # <- in-memory state, lost on restart, wrong across replicas
    log_to_file("app.log", "served " + symbol)   # <- writes to a file, not stdout
    return json_response(price)

app.route("GET", "/health", function(request):
    return json_response({status: "ok"})
)

app.listen(port = 8000)             # <- hardcoded port, ignores PORT

The review pass — check the prompt’s four factors one by one:

  • Factor 3 (config): Failed. API_KEY is hardcoded, and the port is hardcoded too, despite being asked for explicitly.
  • Factor 6 (stateless processes): Failed. CACHE is a plain in-memory dictionary living in the process — it disappears on every restart and gives a different answer on every replica behind a load balancer. This is the single most common thing agents get wrong once you ask for “caching” without also specifying where the cache lives.
  • Factor 7 (port binding): Failed. app.listen(port = 8000) ignores the PORT environment variable entirely.
  • Factor 11 (structured logs to stdout): Failed. It’s writing to a local file, and the message isn’t structured JSON.

The follow-up prompt (this is the realistic workflow — you rarely get all of this right in one shot):

1
2
3
4
5
Fix all four issues:
1. Read API_KEY and PORT from environment variables (fail loudly if missing)
2. Remove the in-memory CACHE dict entirely — for now, call fetch_price on every request
3. Bind to the PORT environment variable, default to 5001 if unset
4. Replace log_to_file with a structured JSON log line written to stdout

Notice the second instruction: rather than asking the agent to “fix the caching,” it’s simpler and safer to remove the premature cache and reintroduce it deliberately later, backed by Redis, once you’ve decided how it should behave across replicas. That’s a judgment call the checklist doesn’t make for you — the 12 factors tell you that something is wrong, not always the least-risky fix.

Worked Example 2 — Reviewing a Generated Circuit Breaker

The review prompt:

1
2
3
4
5
Review this circuit breaker implementation for:
1. Thread safety — does it use a lock before reading/writing state?
2. Correct state transition — does OPEN→HALF_OPEN check if the timeout elapsed?
3. Error propagation — does it re-raise the original exception, not swallow it?
4. Testability — can the state be injected/set for testing without waiting for real timeouts?

Run this prompt against the circuit breaker in the pseudocode above (or a translated, real version of it) as a habit, not a one-off: a circuit breaker is small enough that an agent can produce a plausible one in a single response, and subtle enough that the plausible version frequently fails exactly one of these four checks — most often #4, because “make the timeout configurable for tests” is rarely stated up front and rarely volunteered by the agent unprompted. If a generated breaker hardcodes sleep()-based waiting instead of a comparable, injectable clock, that’s the fix to ask for next.


Pro Tips

  1. Idempotency is not optional. Any operation that can be retried (and they all can be, eventually) must be idempotent. Use idempotency keys on payment and write endpoints.
  2. Health checks have two kinds. A liveness check tells the load balancer the process is running. A readiness check tells it the process is ready to receive traffic (DB connected, cache warm). Implement both — and if you ask an agent for “a health check,” specify which one, since it will otherwise guess.
  3. Timeouts everywhere. Every network call — DB queries, HTTP requests, cache operations — must have an explicit timeout. Never rely on the platform default (often effectively infinite).
  4. The strangler fig pattern for migration. When breaking a monolith into microservices, route new traffic to the new service while the monolith still handles old traffic. Migrate incrementally, not all at once.
  5. Observability trinity: logs, metrics, traces. Logs tell you what happened. Metrics tell you how often and how fast. Traces tell you where time was spent across services. Implement all three from day one.

Common Mistakes

  • Circuit breaker without thread/concurrency safety. Multiple threads or coroutines can read/write state simultaneously. Always use a lock appropriate to your concurrency model.
  • Retry without jitter. If 100 services all retry at the same interval, they create a “thundering herd” that overwhelms the recovering service. Add jitter (randomised delay) to spread retries.
  • Microservices with a shared database. Two microservices sharing one schema are not microservices — they’re a distributed monolith. Each service must own its own data store.
  • Mixing blocking and non-blocking waits in the same runtime. A blocking sleep inside code meant to run on an event loop or async scheduler will stall everything else on that loop — use the async-native wait primitive for your platform.
  • Over-caching. Caching financial data (stock prices) that changes by the second with a 24-hour TTL is worse than no cache — it serves stale data confidently.

References


Next Steps

Continue to Data and AI Architectures →

This post is licensed under CC BY 4.0 by the author.