CODEMOWERS

A field guide for software that runs.

Zen of
Kubernetes.

Build deliberately. Let the platform do its job.

Practical guidance for developing and deploying on bare-metal Kubernetes, also useful on managed cloud and local clusters. Discover the target platform’s capabilities and policies before adapting examples.

  1. Repeating a default in code, an image and a chart makes it unclear which setting controls behavior. Deployment overrides should express actual differences. Read the guidance →
  2. Run the application directly as PID 1, with threads where appropriate. Keep independent process lifecycles in separate containers. Read the guidance →
  3. Keep scheduling, supervision and worker replication in the platform. Application containers should execute work, not run their own orchestration system. Read the guidance →
  4. A replacement Pod must be able to resume service without the old container filesystem. Put durable data in deliberately provisioned backing services. Read the guidance →
  5. Generated secrets and secret stores both work. Keep credentials independent of images and make rotation quick: provision replacements, update every consumer, then revoke old credentials. Read the guidance →
  6. Alert on measurable symptoms with metrics, investigate events with logs, and use traces to locate latency and bottlenecks across service calls. Keep credentials and sensitive payloads out of telemetry. Read the guidance →
  7. Recover only where the failure is understood. Empty results and success acknowledgements must not disguise broken work. Read the guidance →
  8. A worker can finish a side effect and die before acknowledging it. Repeated delivery must not charge twice or duplicate data. Read the guidance →
  9. Readiness controls traffic; liveness triggers restarts. A shared dependency outage should not restart every application replica. Read the guidance →
  10. Use an external identity provider; do not embed an IdP or implement password handling in the application. Keep tenant, record and operation authorization in the application. Read the guidance →
  1. Encryption is useful only with the intended peer. Configure client trust and renewal handling; record the separate Prometheus transport choice. Read the guidance →
  2. An example or platform name does not prove an operator is installed, authorized or healthy. Inspect APIs, controllers and policy bindings. Read the guidance →
  3. Staging and production should run the artifact that was checked. Record image and configuration revisions so a release can be reproduced. Read the guidance →
  4. Stop accepting work on SIGTERM and finish or return in-flight work within the deadline. Correctness must also survive a forced stop. Read the guidance →
  5. Keeping a volume after uninstall does not protect against corruption or data loss. Define backups and verify a recovery procedure. Read the guidance →
  6. The application process must accept both address families; a Kubernetes Service cannot fix an IPv4-only listener. Local Compose may remain IPv4-only. Read the guidance →
  7. Use namespace-local requests for operator or provider-managed databases, caches and identity registrations. Platform administrators manage the shared services. Read the guidance →
  8. Keep platform configuration in discovered, applicable policies. The deployment guide covers networking, TLS, trust, security, storage and ingress conventions, with explicit values when policies do not supply them. Read the guidance →
  9. Treat both as shared infrastructure. Isolate instances with release-specific names and selectors; leave shared operators and namespace lifecycle to the platform. Read the guidance →

From source to service

The guide

The guide

Developing

Back to contents ↑

developing/general

General application development

Markdown ↗

Apply these conventions to new applications. Assess existing applications case by case using the legacy-porting guidance in building/general; do not assume that moving an application to Kubernetes includes rewriting its architecture.

Decisions and configuration

  • Keep one version-controlled codebase per application with many deployments. Do not maintain separate environment-specific source copies; use release configuration for differences. A monorepo can contain multiple applications, and reusable code belongs in declared dependencies.
  • Read existing project decisions before choosing infrastructure or authentication. Ask only for unresolved choices: authentication architecture in developing/architecture, native versus portable dependencies in deploying/storage, and Prometheus transport in deploying/observability. Record each choice, rationale and nonsecret configuration in the project's AGENTS.md or equivalent instructions; preserve unrelated content and update decisions when the user changes them. Resolve conflicts with platform requirements before implementing dependent work.
  • Configure endpoints, credentials and environment-specific behavior at runtime. Treat compatible backing services as replaceable attachments: switching their endpoint or credentials should not require a source change. Use consistent environment-variable names across environments and keep each default in one layer: application, image or deployment. For complex settings, use a configuration library that combines files and environment variables. Helm environment mappings and rollouts belong to deploying/general.
  • Keep code at its sole call site instead of adding single-use helper wrappers. Extract helpers for actual reuse; framework handlers, callbacks and entrypoints are not wrappers.
  • Write readable YAML with indented block mappings and sequences. Retain explicit empty collections where meaningful, and embedded JSON only where the consuming API requires it. End text files with a newline. Ignore generated files and secrets as described in building/general.

Failures and lifecycle

  • Handle only known, recoverable failures at the operation that can recover. Keep try blocks narrow; do not wrap an entire request, job, loop or startup sequence. Avoid blanket exception handlers; where catch syntax is untyped, identify the expected error and rethrow every unrecognized one.
  • Let unexpected failures fail the request, job or process through the runtime/framework. Never turn them into successful acknowledgements, empty results, arbitrary retries or catch-and-log-and-continue behavior. Use cleanup constructs such as finally, defer, context managers or RAII instead of broad catches.
  • Attempt the required database operation directly. Use constraints, atomic operations and specific error handling instead of speculative preflight queries. Fetch required fields once and handle a missing row; do not precede the fetch with an existence check. Business queries using EXISTS remain legitimate. Health endpoints follow developing/health-checks.
  • Start promptly. On SIGTERM, stop accepting requests/jobs and drain, complete or return in-flight work within the termination deadline. Also tolerate abrupt termination: make repeated jobs idempotent or transactional so retries do not duplicate side effects.

Use building/general for the container process model, deploying/workloads for scheduling and resource budgets, and deploying/networking for listeners and TLS. Language sections add runtime-specific details; their frameworks are examples.

developing/architecture

Application architecture and identity

Markdown ↗

Choose the authentication boundary

Before implementing authentication, ask the user to choose a backend for frontend (BFF) or direct OIDC/OAuth API access with custom scopes. Explain the tradeoffs and follow the decision-recording rules in developing/general. Record clients, provider/access proxy, audiences/scopes, token/session ownership and the signed-transfer boundary. Wait for unresolved choices before implementing dependent authentication work.

Pattern Fits Responsibilities
BFF: browser session cookie; backend holds tokens and calls APIs A first-party browser UI; centralized sessions and API composition Secure HttpOnly cookies, CSRF protection and session lifecycle. XSS can still act through the session. Adds a backend hop.
Direct API: clients hold access tokens and call protected APIs Independent browser, mobile, CLI or integration clients; delegated access Audience/scope design, token lifecycle and CORS. Browser tokens are exposed to malicious JavaScript; public clients cannot keep secrets.

Both use an external OIDC provider. Use Authorization Code with PKCE for interactive sign-in. A BFF may also use scoped downstream tokens; token placement and API callers distinguish the patterns. See the IETF browser application guidance.

Recommend based on actual client requirements and obtain the unresolved choice before implementing it. Keep business authorization in the backend; split services only when ownership, scaling or isolation warrants it.

Delegate identity, enforce authorization

  • Use an external identity provider for accounts and authentication. Do not embed an IdP in the application or bundle one into its release; follow deploying/storage for platform ownership and per-application registrations. Do not implement password collection, storage, hashing, verification or reset/recovery flows in the application. Redirect authentication and account lifecycle to the provider; do not add a parallel local login system.
  • Discover provisioning interfaces through deploying/discovery and follow deploying/storage for ownership. Passmower and authentik illustrate OIDC providers; an access proxy such as Pomerium is a separate role, paired with a provider. None is assumed installed.
  • Associate application data with the verified (iss, sub) identity. Enforce tenant membership, object ownership and business permissions; authentication alone does not grant record access.
  • Agree on API audiences and minimal scopes such as documents:read. Verify how the provider grants them: a client's requested scope does not prove a token carries that permission.
  • Validate access tokens with the issuer's supported mechanism, checking issuer, audience, expiry and scopes, then enforce object authorization. ID tokens are not API access tokens. A BFF must enforce the same business permissions rather than proxying broad service credentials without restriction.

Keep signed transfers direct

  • After authorizing an object operation, issue a short-lived, narrowly scoped presigned URL for direct client transfer. Keep storage credentials server-side and configure endpoint CORS for the actual browser origins and operations.
  • Sign for the client-reachable HTTPS endpoint. Preserve the signed method, host, path, query, headers and payload constraints; never rewrite an internal signed URL for public use. See S3 Signature Version 4 query authentication.
  • Do not relay signed-request services such as MinIO/S3 through the BFF or application API. Resolve missing client endpoints instead of adding a proxy workaround. Infrastructure routing must preserve signed requests; other signed services require their own supported direct-client authorization mechanism.

developing/health-checks

Application health checks

Markdown ↗
  • Serve HTTP health/readiness endpoints such as /health and /ready on the internal listener defined in developing/metrics, separate from public application traffic. Wire them into Compose’s healthcheck directive for local use and the corresponding Kubernetes probes. Use a check command available in the image; do not assume a shell or curl exists in a minimal runtime.
  • Keep checks local, cheap and deterministic. Do not query databases or other backing services; a dependency outage must not trigger mass application restarts.
  • Readiness reports whether initialization is complete and the process accepts work; mark it unavailable when draining. Return a success status only while ready and a failure status otherwise.
  • Use a startup probe only when initialization needs a separate allowance before ordinary probes begin. Avoid liveness probes by default; add one only for a demonstrated local failure that restarting can repair. Check process progress, not dependency availability.

developing/languages/go

Developing Go applications

Markdown ↗

Apply developing/general for configuration and lifecycle.

  • Propagate request/job cancellation through contexts and stop background goroutines during shutdown. Bound concurrent work and connection pools.
  • Handle expected errors at the operation that can recover; return unexpected errors with context. Do not ignore errors or use broad panic recovery to keep failed work running.
  • Use one metrics registry across handlers and workers in the process; follow developing/metrics for its listener. Frameworks such as Gin are optional integrations, not deployment requirements.

developing/languages/java

Developing Java applications

Markdown ↗

Apply developing/general for configuration, lifecycle and failures. Choose a framework to fit the application; Spring Boot is one example.

Example: Spring Boot

  • For a Spring Boot HTTP service, include spring-boot-starter-web, annotate the application with @SpringBootApplication, and start it with SpringApplication.run. Use @RestController and @GetMapping to define HTTP handlers.
  • Add spring-boot-starter-actuator and micrometer-registry-prometheus for Prometheus metrics. Keep dependency versions compatible with the chosen Spring Boot release.
  • Use the internal listener required by developing/metrics; expose only required management endpoints. Enable request latency histograms when needed.

For an application listening on port 8080, put this configuration in src/main/resources/application.properties:

server.port=8080
management.server.port=8081
management.endpoints.web.exposure.include=prometheus
management.metrics.distribution.percentiles-histogram.http.server.requests=true

This example exposes /actuator/prometheus on port 8081. Apply developing/metrics for metric safety and deploying/observability for the recorded scrape-transport choice.

developing/languages/nodejs

Developing Node.js applications

Markdown ↗

Apply developing/general for configuration, failures and shutdown, and building/languages/nodejs for the production entrypoint.

  • Keep blocking or CPU-intensive work off the request event loop. Bound outstanding promises, queue consumption and connection pools; an async function does not make synchronous work nonblocking.
  • Await or explicitly supervise asynchronous work so failures reach the request/job owner. Do not consume arbitrary promise rejections with catch-and-continue handlers.
  • Close listeners and drain tracked work on SIGTERM. Choose the framework for application needs and keep health/metrics on the internal listener in developing/metrics.

developing/languages/php

Developing PHP applications

Markdown ↗

Follow building/general for the single-process contract and legacy exceptions. PHP-FPM's master/worker model, even with pm.max_children = 1, does not meet that contract for new services; do not bundle it with nginx and a supervisor. See PHP-FPM configuration.

Example: Framework X

Framework X illustrates a long-lived PHP service using its ReactPHP server and direct php public/index.php entrypoint, rather than FPM integration. Choose a framework to fit the application.

Keep I/O asynchronous throughout the request path; blocking libraries still block the event loop. Use the process's shared metrics registry and bounded shutdown described in developing/metrics and developing/general.

developing/languages/python

Developing Python applications

Markdown ↗

Apply developing/general for configuration, failures and lifecycle, and building/languages/python for packaging.

  • Prefer async I/O with compatible libraries throughout the request path. Do not call blocking database, storage, image-processing or inference operations directly on the event loop. For synchronous applications, retain a single execution thread until concurrency can be introduced correctly.
  • Justify worker threads with measurements and the actual runtime/library behavior. On GIL-enabled runtimes, CPU parallelism requires operations that release the GIL; importing a library alone proves nothing. Bound tasks, thread pools and connection pools; avoid unbounded executors.
  • Follow the single-process contract in building/general; do not add uWSGI worker pools or process supervisors for new services.

Example: Sanic

Sanic is one async HTTP option. Use sanic server:app --host=:: --port=8000 --single-process or app.run(host="::", port=8000, single_process=True) to disable its worker manager and auto-reloader. Its ordinary one-worker mode still includes a manager process. See Running Sanic.

developing/languages/rust

Developing Rust applications

Markdown ↗

Apply developing/general for failure handling and lifecycle. Choose a framework and async runtime for the application's libraries and workload.

  • Use typed errors and propagate failures with context.
  • Keep blocking work out of async request paths. Bound tasks and connections; propagate cancellation and join application tasks during shutdown.
  • Validate required configuration at startup; listener requirements belong to deploying/networking.

Example: Rocket

Rocket applications may use rocket::build() with mounted handlers and configure through Rocket.toml or ROCKET_* variables. Set the deployment listener explicitly rather than inheriting a development bind address.

When using rocket_prometheus, attach instrumentation and export from the same registry, with the internal listener required by developing/metrics. Other frameworks use their equivalent hooks. See building/languages/rust for packaging guidance and building/general for the lolcatz Dockerfile examples.

developing/local-development

Local development

Markdown ↗
  • Use the project's Skaffold or Docker/Compose workflow for development, dependency management, builds and checks. Agents must not install application runtimes, packages or backing services on the host. Reuse existing host orchestration/editing tools; run installation inside build containers. Lockfile-only commands are acceptable when they install no host components.
  • Use Compose for local development and integration testing. Match production's component layout, dependency versions and environment-variable names. Development images may add debugging tools and volume-mounted hot reload.
  • Run Compose without TLS; IPv4-only local networking is acceptable. Keep image and Kubernetes listeners compatible with deploying/networking. Do not repeat defaults already supplied by the application or image.
  • Keep local .env files untracked. Document startup, checks and deployment commands, including required configuration.

developing/logging

Application logging

Markdown ↗

Use logs for triage: record enough context to investigate what happened and why. Use developing/metrics for alerting on measurable symptoms.

  • Disable ANSI color codes and terminal control sequences in Kubernetes logs, including framework and library output. Use the selected logger’s plain-output configuration; do not force terminal colors in containers. Colored output may remain an option for interactive local development.
  • Write to stdout/stderr; let the runtime collect, rotate and retain logs. Do not write container-local log files or duplicate runtime timestamps.
  • Choose plain text or structured JSON for the application's consumers; plain text alone is not a defect. Use meaningful severity levels; exclude credentials and sensitive payloads.
  • Measure repetitive events with metrics instead of producing a log record for every occurrence; see developing/metrics.

developing/metrics

Application metrics

Markdown ↗

Use metrics for alerting: measure symptoms such as failure rates, latency and stalled work, with thresholds tied to actionable operational needs. Use developing/logging for the context needed to triage an alert.

  • Expose Prometheus metrics through the chosen framework's integration or client library. Use one shared, thread-safe registry for all application execution in the container; a scrape must not represent only the worker that answered.
  • Serve metrics and health checks on an internal listener separate from the public application listener. Do not register these routes on the public listener or expose their port through public Ingress or externally exposed Services. Scrape transport and monitoring discovery belong to deploying/observability.
  • Exclude sensitive data from names, labels, values, help text and exemplars: no identities, emails, IP addresses, credentials, tokens, bodies, raw URLs or query strings. Use bounded categories such as route templates, operations and status codes.
  • Measure request latency, request/response sizes, failures and useful business activity such as import throughput. Bound label cardinality. Leave container memory and filesystem usage to runtime monitoring.

The guide

Building

Back to contents ↑

building/general

General container builds

Markdown ↗

Reproducible images

  • Pin base-image versions. Declare all application dependencies and required system libraries explicitly; do not rely on host installations. Commit lockfiles where supported and use locked/frozen installation that fails on manifest disagreement.
  • Install dependencies and compile assets during image builds. Use multiple stages and copy only required runtime files into the final image. Ship a static executable where supported; otherwise include compatible runtime libraries.
  • Keep .dockerignore at the actual build-context root. Exclude Git metadata, host dependencies/build outputs, reports, caches, editor files, local configuration, credentials and unrelated files; retain sources, manifests, lockfiles and required assets. Inspect the context when moving or adding a Dockerfile.
  • Maintain .gitignore separately for generated artifacts and private local files. Ignore rules do not untrack files: inspect accidental tracked outputs before removing them from the index, preserving needed local files. Never copy registry credentials or SSH keys into image layers.
  • Use developing/local-development for build tooling and building/ci for publication and promotion.

Separate executable code from data

  • Package ML model weights, map datasets and similar large static assets in separate data-only OCI images. These assets are inputs to an executable service; do not bake them into the application image or run a container merely to keep the files available.
  • Version application and data images independently, recording compatible digests together in the release. Keep data preparation recipes and immutable source checksums in Git so the artifacts can be reproduced under building/ci. Include only the required data and metadata in a data image; no server, runtime or entrypoint is needed.
  • Mount these artifacts into the consuming workload using the image-volume guidance in deploying/workloads. Separating data does not establish trust in its format or loader; use the application's supported asset formats.

Dockerfile examples

Use the codemowers/lolcatz repository for sample Dockerfiles. The Go, Node.js and Python sections link to published examples. Adapt build contexts, dependencies, toolchains and runtime assets to the application; the examples are not a substitute for the shared build rules.

Runtime contract

  • Run one application process per container, optionally with threads sharing its address space. Launch it directly using JSON exec-form ENTRYPOINT/CMD so it receives SIGTERM as PID 1. If a setup script is unavoidable, end it with exec. Lifecycle behavior belongs to developing/general.
  • Follow deploying/workloads for orchestration. Do not add init wrappers such as dumb-init or tini to new applications. If the application launches subprocesses, it must reap them and propagate cancellation.
  • Supply UID/GID at deployment time, preferably through applicable platform policy. Do not create accounts or bake an application UID into Dockerfile USER. Keep code root-owned and readable/executable by the runtime identity; do not recursively chown it. Use a read-only root filesystem with explicit writable mounts; see deploying/security.

Porting legacy applications

These conventions target newly designed applications. Porting an existing application to Kubernetes is not a promise that it already follows them, nor a requirement to rewrite it before deployment. Evaluate each legacy port case by case rather than treating every deviation as a defect.

  • Inspect the actual process model, identity and filesystem assumptions, state, configuration, networking, startup/shutdown behavior and supported deployment model. Establish what can change without breaking application behavior or vendor support.
  • Agree on the scope of the port. Preserve necessary compatibility, for example an existing master/worker arrangement, startup wrapper, fixed UID or writable directory, when replacing it is outside that scope. Identify which changes are required for the target cluster and which are optional modernization work.
  • Record each retained deviation, its reason, operational consequences and any compensating configuration. Verify startup, health, termination, upgrades and data recovery for the resulting deployment. A documented legacy constraint does not establish that the cluster permits it; resolve applicable policy requirements with the platform owner.
  • Keep exceptions local to that application. Do not turn a legacy workaround into the default for new services or silently expand a port into a redesign.

building/ci

Continuous integration and releases

Markdown ↗

Separate development and production publishing

  • Developers push locally built images only to a separate development registry. Never push or copy a developer workstation's build into the production registry, even when it carries a Git-looking tag.
  • Only trusted CI running from a protected Git branch may publish production images. Restrict production write credentials to that pipeline; developer credentials and untrusted pull-request jobs must have no production push access. Protect the build and workflow definitions along with application code.
  • Make registry/repository prefixes explicit in Helm values and Skaffold environment configuration. Define a chart value such as image.registryPrefix and apply it consistently to application and data-image references; Helm does not supply a standard registry-prefix value. Local development must select the development registry and must never fall back to production.
  • Persist Skaffold's development default-repo in its configuration file with skaffold config set, scoped to a verified development kubecontext. Pass Skaffold's resolved image references into Helm rather than constructing a different destination there. The protected CI pipeline selects the production prefix in its own configuration. See Skaffold image repository handling. Prefixes route builds; registry permissions enforce the boundary.

Example: write a development registry prefix to Skaffold's config, then run against that context. Replace the illustrative context and registry with verified values:

skaffold config set --kube-context verified-dev-context default-repo registry.dev.example.com/team
skaffold dev --kube-context verified-dev-context

This is Skaffold's per-user configuration (normally ~/.skaffold/config), not a default-repo field in skaffold.yaml. Avoid a global production default on developer machines.

Reproducible, immutable production artifacts

  • Every production image must be reproducible from its recorded Git commit and declared inputs. Commit Dockerfiles, build configuration and lockfiles; pin base images and external build inputs by digest or checksum, and retain them. Never depend on uncommitted files, workstation outputs or unversioned downloads. Large model/data inputs may live outside Git, but their immutable references, checksums and preparation instructions belong in Git.
  • Automate builds, checks and vulnerability scanning in CI. Run integration checks through Compose with production-aligned dependencies, and verify the packaged image in Kubernetes. Build once per source revision and promote that verified CI artifact through staging and production without rebuilding or replacing it with a local build.
  • Attach the applicable OCI image annotations to every built image, including its source URL, revision, creation time, version, title and description. Set license metadata only when the project's license is established. Derive values from the same immutable source revision and release metadata used for the image tag.
  • Enforce immutable production tags in the registry and deploy by image digest. Never overwrite a published production tag. Mutable tags such as latest belong only to development publishing/defaults; release values must resolve to immutable references. Apply the same rules to separate data images.
  • Record source, application/data-image digests and deployment configuration for each release. Retain artifacts needed to redeploy or roll back. Supply environment configuration at deployment time; check schema compatibility before rollback and coordinate migrations as described in deploying/workloads.

building/languages/go

Building Go containers

Markdown ↗

Apply building/general for image and runtime conventions.

  • Copy go.mod and go.sum before sources; run go mod download to cache dependency downloads separately.
  • Use CGO_ENABLED=0 only when dependencies support it. CGo applications need compatible runtime libraries or verified static linking. Optional -ldflags="-s -w" removes symbol/debug information.
  • Copy assets to their runtime paths. A scratch image also needs any CA bundle or other runtime data used by the application; verify the packaged image starts and serves required assets.

See the lolcatz Go service Dockerfile for a complete example.

building/languages/java

Building Java containers

Markdown ↗

Use the project's build tool and a compatible JVM runtime image; apply building/general for shared image conventions. A JVM application cannot run directly in scratch.

  • Declare the Java version in build configuration and align builder/runtime versions. Choose maintained distributions, for example Temurin or Corretto.
  • For Maven with Spring Boot, use spring-boot-maven-plugin to produce an executable JAR. Keep build-time checks enabled and match the copied filename to the actual artifact.

See building/general for the lolcatz Dockerfile examples; adapt the build and runtime stages to this language’s toolchain.

building/languages/nodejs

Building Node.js containers

Markdown ↗

Apply building/general for image stages, locked dependencies, runtime identity and process ownership.

  • Use the same pinned Node version across stages. Omit development-only packages from the runtime image.
  • Launch Node directly with exec-form CMD ["node", "server.js"], using the actual built entrypoint. Avoid production startup through npm start, npm run or npx: wrappers add process layers and may run hooks or resolve packages.
  • Package-manager scripts remain suitable for installation, builds, checks and local development. A start script is not a defect when production launches Node directly.

See the lolcatz Next.js frontend Dockerfile for a complete example.

building/languages/python

Building Python containers

Markdown ↗

Apply building/general for locked dependencies, runtime libraries and image contents. Install into the container's Python environment without an extra virtual environment; with uv, set UV_PROJECT_ENVIRONMENT=/usr/local.

Virtual environments remain appropriate outside containers. Agent development workflow follows developing/local-development.

See the lolcatz Python OCR service Dockerfile for a complete example.

building/languages/rust

Building Rust containers

Markdown ↗

Apply building/general for image stages, runtime identity and file permissions.

  • Cache dependency work using Cargo manifests and the committed lockfile; include required workspace member manifests. Use cargo fetch --locked and cargo build --release --locked --target <target>. Do not update dependencies during image builds.
  • If caching builds with placeholder sources, replace them with the complete real source tree and ensure Cargo rebuilds it. Exclude host target artifacts.
  • Choose the target and linker for the deployment architecture. Native dependencies also need compatible libraries. Verify there is no dynamic interpreter or required shared library before choosing scratch.
  • Copy required assets, including CA trust data when needed. Use a compatible minimal runtime if static linking is unsuitable. Verify startup and TLS/native-library paths in the final image.

See building/general for the lolcatz Dockerfile examples; adapt the build and runtime stages to this language’s toolchain.

The guide

Deploying

Back to contents ↑

deploying/general

General deployment guidance

Markdown ↗

Discover the target using deploying/discovery, select dependencies using deploying/storage, and deploy immutable releases as described in building/ci. Keep configuration in version control without credentials. Configure Helm image prefixes and Skaffold publishing destinations according to building/ci. YAML formatting and project decision records follow developing/general.

Release ownership and isolation

  • Namespace lifecycle belongs to the platform. Never add kind: Namespace to application charts, manifests or GitOps inventories, or adopt an existing namespace into application ownership.
  • Before removing an existing Namespace declaration, inspect live ownership/tracking metadata and Helm/GitOps release inventories. Detach it safely from application ownership and pruning first; labels alone do not remove inventory entries. If safe detachment is unresolved, retain the declaration until the controller/release can be migrated. Never delete/recreate the namespace.
  • Support independent instances across namespaces and multiple Helm releases within one namespace. Derive application-owned names and references from .Release.Name using consistent chart helpers. Include release identity in Pod labels and workload/Service selectors. Cover dependency requests, identity registrations, Secrets, ConfigMaps and PVCs so releases cannot collide or select each other's Pods.
  • Omit ordinary metadata.namespace; let the release target namespace apply. Use .Release.Namespace only for required explicit references, such as RBAC ServiceAccount subjects or namespace-qualified addresses. Cluster-scoped resources never have a namespace. Configure instance-specific hostnames through values.

Helm configuration and admission defaults

  • Expose TLS and other platform settings through Helm values, but leave their defaults unset. This includes certificate/issuer references, trust, listeners, storage classes, replicas and security settings as appropriate. Render fields only when explicitly supplied; do not reproduce platform defaults with Helm's default function.
  • Discover applicable mutation policies and resource defaults before relying on them. Kubernetes 1.36 enables the stable MutatingAdmissionPolicy mechanism; it does not install platform policies or bindings. Supply required settings through deployment-specific values when no applicable policy or resource default provides them. Keep resource-specific conventions in their own cluster descriptions.
  • Applicable admission policies may configure IPv6 listeners for databases and other dependencies, database TLS, CA-certificate mounts and trust paths, Pod/container security contexts, dependency storage classes, and ingress conventions. Discover each policy’s target resources, bindings and requirements independently; none of these capabilities implies the others. Keep concrete defaults in the responsible policy/resource description. Verify the admitted configuration and the operator/application’s resulting behavior; a mounted CA bundle still needs the client to use it. See deploying/networking, deploying/security and deploying/storage for the corresponding requirements.
  • Unset means omitted, not disabled or empty. Do not render nulls, empty strings or empty collections as placeholders. Check key presence to preserve explicit false and 0; retain empty collections only where the API requires them. Example values are not chart defaults.
  • Detect optional APIs with Helm capability checks, separately for each kind; do not use a feature flag as proof that infrastructure exists. Monitoring integration follows deploying/observability; required dependencies must be resolved rather than silently omitted.
  • Render the chart, inspect a server-side dry run where supported, then verify admitted resources and runtime readiness. Omission alone does not prove the platform supplied a setting.

Environment mappings and rollouts

  • Generate an explicit env: entry for each configured variable; do not use envFrom. Render nonsecret values as quoted value: strings in the Pod template. Use explicit secretKeyRef or configMapKeyRef mappings for referenced data; never inline generated credentials into Helm values or manifests.
  • A changed Pod template triggers the Deployment's rollout. Changes to referenced Secret/ConfigMap data alone do not: use a configuration checksum annotation for chart-owned data, or the established reload/restart mechanism for operator-owned data.
  • Apply unset semantics to platform environment settings too. Omit unspecified variables and omit env when there are none. Preserve explicitly supplied booleans and numbers as strings.

Example rendered entries for explicitly supplied values:

env:
  - name: LOG_LEVEL
    value: "debug"
  - name: TLS_ENABLED
    value: "true"

deploying/discovery

Discover the target cluster before provisioning

Markdown ↗

Identify the target and inspect its capabilities before provisioning. Use advertised Driftmower tools, direct kubectl, or both when a response lacks detail. CRD-specific examples belong to the target cluster advisory; examples never establish availability.

With Driftmower MCP

Identify the connection's cluster using get_cluster_identity and its platform using get_cluster_platform when advertised. Record whether it is k3s, vanilla Kubernetes, EKS, AKS or another verified platform. Treat unknown or conflicting evidence as a question to resolve with the user, not as vanilla Kubernetes. Before using local kubectl, match its Kubernetes API URL to a verified kubeconfig context. Discover the tools advertised by that connection instead of assuming a particular Driftmower version exposes every tool.

Information Advertised tools to use
Platform, Kubernetes version and detection evidence get_cluster_platform
Installed CRDs, served versions, ingress classes and issuers describe_cluster_capabilities
A particular CRD's live schema and description describe_custom_resource_definition with its full CRD name
Admission policies, descriptions, target rules and binding scope list_admission_policies
StorageClass descriptions, parameters, defaults and lifecycle settings list_storage_classes
Accessible namespaces and local guidance list_accessible_namespaces, get_cluster_advisory

Use additional offerings or namespace-readiness tools only if the connection actually advertises them. A capability listing does not prove controller health, permission to create a resource, or that a policy applies. The admission catalog is a discovery summary, not an evaluation of policy code: inspect matching policy and binding specifications, selectors, conditions and parameters with kubectl when the summary is insufficient. Missing or inaccessible information is unknown, not evidence that a policy or resource is absent.

Without Driftmower, or when more detail is needed

Use the existing kubeconfig credentials and an explicit verified context on every cluster operation. Confirm the target API URL with the user or established project configuration if no MCP identity is available. Do not switch the current context implicitly or print raw kubeconfig credentials. Choose an existing target namespace; checking it does not transfer namespace lifecycle ownership to the application.

Replace verified-context and existing-namespace in each command with the verified deployment target:

kubectl --context verified-context api-resources
kubectl --context verified-context get namespace existing-namespace -o yaml
kubectl --context verified-context get storageclasses.storage.k8s.io -o yaml
kubectl --context verified-context --namespace existing-namespace get resourcequotas,limitranges -o yaml

Read annotations and descriptions as well as settings. For StorageClasses, inspect the provisioner, default annotation, parameters, reclaim policy, binding mode, expansion support and topology restrictions. Do not assume a default StorageClass exists or copy one from a different cluster.

Determine platform and record infrastructure choices

Read the server version and, when permitted, platform-specific node labels:

kubectl --context verified-context version -o yaml
kubectl --context verified-context get nodes -L eks.amazonaws.com/nodegroup,kubernetes.azure.com/cluster

A k3s version suffix, EKS distribution markers or AKS-specific labels are evidence to check against the cluster's provisioning records. A generic version, cloud provider ID or missing label does not prove vanilla Kubernetes or a particular managed service. If discovery is restricted or ambiguous, confirm the platform with the user or administrator. Do not infer the platform from an MCP registration name or kubeconfig context name.

Resolve dependency and Prometheus choices through deploying/storage and deploying/observability, using the decision-recording rules in developing/general.

Check each required CRD

For each selected operator API, including native-service provisioning APIs, substitute its CRD, resource name and served version:

CRD_NAME=resourceplural.operator.example
RESOURCE_NAME=resourceplural
API_VERSION=operator.example/v1
kubectl --context verified-context get customresourcedefinitions.apiextensions.k8s.io "$CRD_NAME" -o yaml
kubectl --context verified-context explain "$RESOURCE_NAME.spec" --api-version="$API_VERSION"
kubectl --context verified-context --namespace existing-namespace auth can-i create "$CRD_NAME"

Inspect the CRD's Established condition, deletion state, served versions, scope and OpenAPI schema. Use the resource's discovered API version for kubectl explain. Repeat for every other custom resource in the chosen example, including certificate and credential generators when referenced. A CRD alone does not establish that its controller is running, watches the target namespace or can reconcile the resource: inspect the discovered controller workload and relevant reconciliation status with the permissions available.

If the API or its controller is absent, establish how that prerequisite will be provided before deploying dependent resources. Do not silently install an operator, copy CRD definitions into an application chart or redirect to a named alternative whose availability has not been established. Distinguish Forbidden, discovery/network failures and actual NotFound results.

Check admission configuration and bindings

Discover which admission APIs are served first:

kubectl --context verified-context api-resources --api-group=admissionregistration.k8s.io

Run the corresponding reads for the APIs present in that result:

kubectl --context verified-context get mutatingadmissionpolicies.admissionregistration.k8s.io -o yaml
kubectl --context verified-context get mutatingadmissionpolicybindings.admissionregistration.k8s.io -o yaml
kubectl --context verified-context get validatingadmissionpolicies.admissionregistration.k8s.io -o yaml
kubectl --context verified-context get validatingadmissionpolicybindings.admissionregistration.k8s.io -o yaml
kubectl --context verified-context get mutatingwebhookconfigurations.admissionregistration.k8s.io -o yaml
kubectl --context verified-context get validatingwebhookconfigurations.admissionregistration.k8s.io -o yaml

Match policies to bindings, resource rules, operations, namespace/object selectors, conditions and referenced parameter resources. Check namespace labels and validation actions; an unbound policy or a nonmatching binding does not supply defaults. Webhook configurations may establish another admission mechanism; inspect its relevant configuration and documented behavior when present. An empty API listing does not rule out other admission configuration controlled by the cluster administrator. If inspection is forbidden or incomplete, obtain the missing requirements rather than assuming there are no constraints.

Check TLS and other prerequisites

Discover issuer and certificate APIs before querying their instances. If cert-manager's APIs are installed, inspect the available cluster-scoped and namespaced issuers:

kubectl --context verified-context get clusterissuers.cert-manager.io -o yaml
kubectl --context verified-context --namespace existing-namespace get issuers.cert-manager.io -o yaml

Read each issuer's own description and readiness to determine its suitability. Discover the trust bundle and the namespace's network posture from relevant ConfigMaps, NetworkPolicies and platform guidance. Inspect only the required objects and Secret key names; do not print credential values. No issuer, trust-bundle name or secret generator is implied by another resource's example.

After discovery, apply the Helm defaults and deployment verification rules in deploying/general.

References: Kubernetes CRDs and schema discovery, admission policies and bindings, and admission webhooks.

deploying/networking

Deployment networking

Markdown ↗
  • Use service names in Compose and namespace-qualified service names in Kubernetes. Align Service selectors, target ports and application listeners; configure instance hostnames through deployment values.
  • Support IPv6 and IPv4 in Kubernetes: bind to :: with IPv4 acceptance or provide separate listeners. 0.0.0.0 alone is IPv4-only; a Service cannot change the process's bind address. Inspect image defaults and runtime overrides together. Local Compose may use IPv4-only bindings.
  • Bound inter-service calls with timeouts; retry only operations that tolerate repetition, using backoff and a finite budget.
  • Keep TLS configurable in software. Local Compose runs without TLS; Kubernetes application and backing-service connections use verified TLS. Prometheus transport is a separate choice in deploying/observability. Never bake deployment certificates or trust paths into images.
  • Use deploying/discovery to find suitable issuers, allowed hostnames and trust resources. Prefer an available issuer whose documented purpose meets the connection's requirements; provision a new one only when necessary. Helm unset/default behavior belongs to deploying/general.
  • Configure servers to load certificates and clients to trust the selected issuer and verify hostnames. Mount the discovered trust bundle when needed; certificate issuance does not configure client trust. Arrange reload or restart after renewal. TLS does not replace application authentication.
  • Expose browser-facing routes through Ingress. Align certificate DNS names, Ingress TLS hosts, public URLs and OIDC redirect URIs. Keep background workers and backing-service administration internal; health and metrics use the listener in developing/metrics.

deploying/observability

Deployment observability

Markdown ↗

Use metrics for alerting, logs for triage and traces for performance tuning. Traces help locate latency and bottlenecks across service calls; correlate them with logs and metrics when investigating a slow operation. Use the platform’s available tracing pipeline, keep collection configurable, and exclude credentials and sensitive payloads.

Application instrumentation belongs to developing/metrics, developing/health-checks and developing/logging.

Prometheus transport decision

Ask the user to choose internal HTTP or HTTPS with TLS before configuring scraping. Follow the decision-recording rules in developing/general. Explain that HTTP is unencrypted, while HTTPS requires certificates, verified CA/hostname trust and renewal/reload handling. Inspect policy constraints before presenting the choice; resolve conflicts with an existing decision rather than silently changing it.

Record the transport and, for HTTPS, its nonsecret certificate/trust provisioning approach. Do not infer metrics TLS from application TLS or fall back to HTTP when certificates are missing. Keep listener and scrape settings consistent with the decision; both modes remain internal.

Monitoring integration

  • Discover the installed monitoring APIs and controller. Use its cluster-specific resource guidance; without an operator, use the established Prometheus configuration mechanism. Do not install monitoring infrastructure through the application chart.
  • Detect optional resources independently with Helm capabilities, for example monitoring.coreos.com/v1/PodMonitor and monitoring.coreos.com/v1/PrometheusRule. Match release-specific workload labels and the monitoring system's discovery selectors. Keep application metrics available when these APIs are absent.
  • Follow the unset Helm defaults in deploying/general; provide explicit deployment values for required settings that policies or resource defaults do not supply. Inspect admitted configuration and verify scraping succeeds.
  • Configure health probes on the internal listener and collect stdout/stderr through the platform log pipeline.

deploying/security

Runtime security

Markdown ↗
  • Run workloads as a non-root identity supplied by applicable platform policy or explicit deployment values; do not assume that omitting image USER ensures non-root execution. Use a read-only root filesystem with explicit writable mounts. Keep runtime identity and Pod security configurable with unset chart defaults; follow building/general for image permissions and deploying/general for admission behavior. If no applicable policy provides seccomp, supply seccompProfile.type: RuntimeDefault at Pod level in deployment values.
  • Keep credentials out of source control, images and logs. Use the platform’s secret generator, secret store or other discovered provisioning mechanism; commit declarations, never generated values. Consume Kubernetes Secrets through explicit references or mounted files. Locally, use untracked configuration or Docker secrets.
  • Make rotation fast and repeatable without rebuilding images. Give separate component relationships their own credentials; reuse operator-generated credentials when the operator owns authentication. Provision replacements, update all consumers through their supported reload or rollout mechanism, verify access, then revoke old credentials. Allow an overlap where supported for routine rotation; compromised credentials may require immediate revocation. Keep the rotation procedure documented and verify it works. Network trust requirements belong to deploying/networking.
  • Provision OIDC registrations through the selected identity service, following developing/architecture and the dependency ownership rules in deploying/storage. Align issuer, audience and HTTPS redirect URIs. Keep confidential client credentials on the server.
  • Provision session-signing secrets once and share them across application replicas. Rotate deliberately; do not regenerate them on every deployment.

deploying/storage

Application dependencies and persistence

Markdown ↗

Native or portable services

Determine the platform using deploying/discovery, then ask the user to choose provider-native managed services or portable Kubernetes operators. Record the platform, approach, selected service/operator and rationale using developing/general; reuse existing decisions.

For PostgreSQL on AWS, Amazon RDS versus CloudNativePG illustrates the tradeoff: provider operations versus a reusable Kubernetes deployment model with operator, storage, backup and recovery responsibilities. Check availability, permissions, connectivity, costs and support before recommending either. A platform name does not establish service availability. Resolve a missing prerequisite with the user/platform administrator rather than silently changing the selected approach.

Ownership and provisioning

  • Prefer custom resources for application databases, caches and OpenID Connect registrations. Passmower, CloudNativePG (CNPG) and Dragonfly are examples, subject to deploying/discovery. For native services, prefer an available Kubernetes provisioning interface such as Crossplane; verify its providers, credentials, permissions and resource scope.
  • The platform team administers databases, caches and identity services cluster-wide, including operators, access policy, upgrades, backups and recovery. Application releases own per-instance requests and registrations. Never install shared operators, identity infrastructure or external dependency charts through the application chart. Do not handwrite StatefulSets, Deployments or Pods for these backing services; operator-generated workloads stay under operator management.
  • Treat the cluster as shared. Keep per-instance dependency requests, identity registrations and credential references in the application's namespace, even when the requested service runs remotely. For a cluster-scoped API, use an available namespaced request interface or administrator-provisioned resources with scoped connection information.
  • Apply the release naming rules in deploying/general to support multiple instances in one namespace. Share a service only through an explicit recorded choice, with isolated databases, cache access and identity registrations. Credential lifecycle belongs to deploying/security.

Data lifecycle

  • Keep application containers stateless. Store session state needed across requests in an appropriate shared backing service or supported client-side session design; do not rely on a replica’s memory/filesystem or sticky routing for correctness. Choose capacity, storage class, retention, backups and restore procedures for the workload; verify recovery. A retained PVC or Helm keep annotation is not a backup. Use bounded disposable volumes for temporary files.
  • Configure internal and browser-facing object-storage endpoints separately. Follow developing/architecture for direct signed transfers and deploying/networking for TLS.
  • Obtain custom-resource schemas and examples from the target cluster's Driftmower advisory or installed operator documentation. This general guide does not define operator defaults.

deploying/workloads

Kubernetes workload design

Markdown ↗
  • Separate HTTP servers and background workers so their replicas and resource budgets scale independently. Reuse an image when they share code, selecting distinct startup commands. Workers without inbound network traffic need no Service. Process and shutdown rules belong to building/general and developing/general.
  • Do not build an orchestrator inside an application container: no forked worker managers, supervisord/s6 process trees, embedded message brokers used to dispatch child workers, or custom restart/replication loops. Use Kubernetes workloads for process lifecycle and independently provisioned messaging services when durable queues are needed. An application may consume a queue and schedule bounded in-process tasks; it must not supervise a fleet of child services. The narrowly scoped legacy exception is described in building/general.
  • Use Kubernetes Lease objects for leader election, CronJobs for schedules with minute-level resolution, and Jobs for finite work. Keep scheduling and supervision out of long-lived application containers.
  • Package migrations and maintenance commands in the application image and run them as Jobs with the release's configuration. Coordinate execution so replicas cannot race migrations; support overlapping application versions during rollouts.
  • Set termination grace periods to cover bounded request draining and job completion/return. Set job retry and concurrency limits for the operation's failure and duplicate-delivery semantics.
  • Size CPU and memory per component, including heavier media workers. Inspect namespace quotas and defaults using deploying/discovery; supply values required by the workload under deploying/general.
  • Discover messaging APIs before declaring resources. Choose partitions, retention, replay and deletion behavior deliberately. Compacted topics represent keyed state and do not preserve every historical event.

Mount packaged data with image volumes

For the separate model/map/data images in building/general, prefer Kubernetes image volumes when supported by the target cluster and container runtime. Mount the artifact's content into the application as read-only data; do not start a sidecar just to hold files. Keep mutable outputs on separate writable storage.

Expose the data-image reference and mount path through deployment configuration, pin production artifacts by digest, and verify registry access and the required image-volume support before use. Follow the official image-volume walkthrough. If support is unavailable, agree on an available data-delivery mechanism while retaining separate, versioned artifacts; do not silently bundle the data back into the executable image.

For your coding agent

The same guide, over MCP.

Connect using Streamable HTTP. No sign-in or API key is required. Choose project scope to share setup with a repository, or user scope to use it across projects.

https://mcp.codemowers.io/mcp

Codex

Global · your user account

This command adds the server to your user configuration, normally ~/.codex/config.toml:

codex mcp add codemowers --url https://mcp.codemowers.io/mcp

Project · this repository

Add this table to .codex/config.toml in the project root, preserving existing settings. Codex loads project configuration for trusted projects.

[mcp_servers.codemowers]
url = "https://mcp.codemowers.io/mcp"

Codex MCP documentation ↗

Claude Code

Global · your user account

claude mcp add --transport http --scope user codemowers https://mcp.codemowers.io/mcp

Project · this repository

Run from the project root. This creates or updates .mcp.json, which can be committed for the team. Claude Code prompts before using project-scoped servers.

claude mcp add --transport http --scope project codemowers https://mcp.codemowers.io/mcp

Claude Code MCP documentation ↗

Run locally with Docker

Once an image is published, replace VERSION with its published tag, or use an image digest:

docker run --rm --name codemowers-mcp \
  -p 127.0.0.1:3000:3000 \
  -e PUBLIC_URL=http://localhost:3000 \
  ghcr.io/codemowers/codemowers-mcp:VERSION

Open localhost:3000 ↗ for the guide. For local MCP access, substitute http://localhost:3000/mcp in the setup snippets above. Stop the container with Ctrl-C.

Call get_engineering_advisory without arguments for the index, or supply a document path. Documents are also available as MCP resources. Discover the target cluster’s capabilities and policies before provisioning.