Platform Engineering

The AWS spine, built by engineers who run it.

We install the multi-account AWS Organization, harden it against the controls a real production account needs, ship the ECS pattern your app runs on, and hand you the Terragrunt repo. Our own production SaaS apps run on this pattern today. The IaC is yours from day one.

What we ship you is what we run. Hardened by years of production, not a weekend of vibes.

30 minutes. 1-page PDF in 2 business days.

What Platform Engineering means here.

Three layers, delivered as one engagement or as separate steps. Buy one, buy the pair, or bundle all three.

The foundation

A multi-account AWS Organization with a delegated Audit account, a Log Archive account, IAM Identity Center with least-privilege Permission Sets, a 3-tier VPC with private access via Twingate, org-wide CloudTrail, and AWS Config. Everything as Terragrunt code, in the module-and-live split we run on our own accounts.

  • AWS Organizations
  • IAM Identity Center
  • Twingate
  • CloudTrail
  • AWS Config

The hardening

The preventive controls layer that stops a CIS/FSBP-flaggable resource from ever being created — CIS/FSBP-targeted SCPs, AWS Config auto-remediations for the safe-to-auto-fix patterns and detect-and-alert with a runbook for the rest, Security Hub with the standards enabled and every FSBP finding triaged, GuardDuty across the org, IAM Access Analyzer with cross-account findings routed, and customer-managed KMS with explicit rotation.

  • Security Hub
  • GuardDuty
  • Access Analyzer
  • SCPs
  • AWS Config

The application platform

A shared ECS Fargate substrate — cluster, shared ALB with host-based routing, ECR registry policies, account-level alarms — and a per-app deployment surface (ECS service, ECR repo, migrations task, ACM/Route53, Secrets Manager, per-app KMS, an ARN-scoped GitHub OIDC deploy role, S3 buckets, and the alarm pack that catches what AWS-native monitoring can't see).

  • ECS Fargate
  • ALB
  • ECR
  • GitHub OIDC
  • CloudWatch

What's the thing you most need handled?

One question. Three answers. Each routes to the specific deliverables list below — not a sales pitch, the actual scope.

Start with the foundation

Start with just the foundation.

Multi-account Organization, Identity Center, 3-tier VPC + Twingate, CloudTrail, AWS Config, IaC handover. No preventive controls layer — that's Core+. For teams who want the substrate done right before the hardening conversation.

Core 4 weeks

Read the Core scope
Harden Preferred

Get our AWS hardened the right way.

Preventive controls, customer-managed KMS, Security Hub + FSBP, GuardDuty, Access Analyzer, and the engineering platform (Terraformed GitHub + centralized AI skills + onboarding CLI) bundled together. Built for engineering teams, not compliance funnels.

Core+ 6 weeks

Read the Core+ scope
Ship

Ship this app on AWS in weeks, not quarters.

The shared-ALB ECS Fargate pattern that runs our own SaaS apps, with ARN-scoped OIDC deploys and the log-metric-filter alarm pack you keep. Three variants — new AWS account, existing substrate, or an additional app on a platform we installed.

Container Platform — ECS 4–6 weeks

Read the Container Platform — ECS scope

What ships in each engagement.

Prices are individually anchored. Combinations are negotiated per deal — there are no pre-priced bundles. The Terragrunt repository and every line of code is yours from day one.

P1

Core AWS Production Foundation

4 weeks fixed price, fixed scope

The multi-account foundation, done right, before the hardening conversation.

If you're pre-hardening, pre-audit, and want the substrate laid on patterns we've lived with across our own production accounts, this is where you start.

What ships

AWS Organization + Audit / Log Archive accounts + Delegated Admins

  • AWS Organization with a delegated Audit account and a Log Archive account — additional accounts created per your project or environment split, not a fixed template. Control Tower optional.
  • Delegated administrators for the services that support delegation (Security Hub, GuardDuty, Config, Access Analyzer). The management account stays out of the day-to-day.
  • Governance SCPs and Tag Policies attached at the org root (deny-root, region restrictions, account-deletion protection, enforced tag vocabularies).

Base SCPs + Permission Sets + Billing Alerts

  • AWS IAM Identity Center (SSO) with custom Permission Sets — least-privilege groups mapped per access pattern; users assigned to the right group.
  • GitHub OIDC trust wired to the CI baseline (ARN-scoped, no long-lived AWS keys in CI).
  • IAM password policy enforced by code — 14-character minimum, 90-day max age, 24-character reuse prevention.
  • Billing alerts — AWS Budgets thresholds per account, Cost Anomaly Detection alerts routed to Slack or email, untagged-resource detection.

3-tier VPC + NAT Gateway + VPC Endpoints + VPN

  • 3-tier VPC (public / private / data subnets), NAT gateway, security-group baselines, Route53 hygiene.
  • VPC Flow Logs at INFREQUENT_ACCESS (full audit trail at ~1/3 the cost of standard CloudWatch Logs).
  • Nine VPC endpoint types templated (S3, ECR, Secrets Manager, SSM, KMS, and more) — private-subnet workloads talk to AWS over PrivateLink instead of NAT.
  • Twingate VPN connector (or Tailscale / AWS Client VPN) — private access to the data tier and admin surfaces, no bastion hosts, no public database endpoints.

CloudTrail centralized + Log Archive centralized + Athena log queries

  • CloudTrail centralized — org-wide, multi-region, delivered to the Log Archive account. Auto-enrolls every future account you add.
  • Log Archive account with a restrictive bucket policy — cross-account read for the SOC or auditor; no cross-account write.
  • Athena workgroup wired against the CloudTrail bucket — searchable audit trail on day one, no ad-hoc S3 grep archaeology when an incident starts.

AWS Config (org-wide) + cost-tuned recording

  • AWS Config enabled and recording across every account in the Organization, delivered to the Log Archive account.
  • Cost-tuned recording — only the resource types Security Hub / FSBP actually reads, not the full firehose. AWS Config bills stay predictable.

Everything coded — full source repository delivered to your GitHub org

  • Module-and-live IaC split — every AWS resource lives in a versioned, semver-tagged module in a centralized platform-terraform-modules repo, consumed from a thin live-config repo. Independent per-module versioning, auto-tagged on merge.
  • State backend hardened by default — S3 with encryption at rest, DynamoDB lock table, and state access via a separate cross-account assume-role. State theft requires two compromises, not one.
  • Pre-commit gate chain — terraform_fmt, terraform_tflint, terraform_trivy IaC vulnerability scanning, terraform_validate, plus ticket-link enforcement. No commit lands without the chain.
  • GitHub Actions TerragruntDeployment workflow with hard approval gate — plan on every dispatch, apply only after a GitHub Environment reviewer approves. OIDC into the deploy role, no long-lived AWS keys.
  • Provider default_tags discipline — every resource auto-tagged with project_name, project_short_name, environment, component, managed_by=Terraform. Cost allocation works on day one.
  • Runbook + architecture diagram + onboarding doc for the next engineer, plus a 90-minute recorded IaC walkthrough.
Read the full Core scope
Preferred tier

P2

Core+ Foundation + Security Hardening + Engineering Platform

6 weeks fixed price, fixed scope

The preferred tier. Everything in Core, plus preventive controls that stop a CIS/FSBP violation from being created in the first place.

Sold as one engagement because the engineering leader who needs the hardening is almost always the same one who feels the developer-platform pain — and pricing them as separate SKUs created shelf-shopping friction without revenue upside.

What ships

Everything in Core, plus…

  • The full account structure, identity model, network, logging, and IaC handover from the Core engagement. Core+ is Core plus four additional tracks — Security Hub + FSBP/CIS, detective services, preventive SCPs, and the engineering platform — delivered in the same 6-week window.

Security Hub + AWS FSBP/CIS conformance packs

  • AWS Foundational Security Best Practices (FSBP) — Security Hub standard enabled across every account, every FSBP finding triaged. High/critical remediated in code; medium documented with risk acceptance or a remediation plan.
  • CIS AWS Foundations Benchmark — every applicable control (Benchmark version pinned at delivery time) implemented in code; non-applicable controls explicitly documented with rationale.
  • Conformance-pack coverage report — a rendered summary of which controls pass, which are accepted-risk, and why. The document a SOC2 auditor asks for.

GuardDuty + Inspector + IAM Access Analyzer

  • GuardDuty across every account and region — findings routed to Slack or email.
  • AWS Inspector enabled for EC2 / ECR / Lambda — CVE surface visible at the image level, not just at runtime.
  • IAM Access Analyzer with cross-account findings routed — external-access surprises land in the queue, not in an audit six months later.
  • AWS Macie for the prod data account (optional, your call).

Security SCPs + Guardrails SCPs + AWS Config auto-remediations

  • CIS/FSBP-targeted Service Control Policies — the exact actions that would create a non-compliant resource are denied at the org root, so violations don't reach the Security Hub queue in the first place.
  • Guardrails SCPs — deny public S3, deny public snapshots, deny IAM user creation, deny IMDSv1, deny root-user actions outside the management account. The rails an engineer can safely work inside.
  • AWS Config auto-remediations for the safe-to-auto-fix patterns. Controls where auto-remediation has too much blast radius ship as detect-and-alert with a runbook — we're honest about the difference.

Terraformed GitHub + agent-agnostic AI skills + onboarding CLI

  • Terraformed GitHub — organization configuration as code (security policies, SSO enforcement, member roles, IP allowlists). A reusable repository module wires branch protections, secret scanning + push protection, Dependabot, and OIDC trust to AWS on every new repo. Up to 10 existing repositories imported and standardized (additional repos priced separately).
  • Centralized reusable GitHub Actions workflow library — OIDC-auth, build, test, lint, security-scan, ECR publish, ECS deploy, CloudFront deploy — versioned as a shared library and consumed by tag. Updates propagate by version bump, not copy-paste.
  • Pre-commit + commit-message chain shipped to every new repo — ticket-link enforcement ((PROJ-NNNN) prefix + linked-issue body line), Biome format + lint, type-check, build. Wires straight into your Jira/Linear project.
  • Centralized, agent-agnostic AI skills repository — one curated repo of skills, prompts, and tool allowlists that every engineer consumes from, usable with Claude Code, Cursor, and any agent that accepts a system prompt + tool allowlist. Ships with a set of starter skills — cross-account AWS read-only inspection, Terragrunt plan/apply guidance, CI-workflow debugging, knowledge-base entry, and the centralized allowlist itself. Client owns the repo; extend or fork freely.
  • Allowlist registry — allowed-commands.yaml compiled into Claude Code permissions, Cursor allow-lists, and any agent's tool-policy format via a single install script.
  • Developer-onboarding CLI — a one-command bootstrap for a new engineer that installs your agent(s) of choice, wires the centralized skills repo, configures AWS SSO + role-assumption profiles for every account, installs required dev tools via brew or asdf, sets up git identity + commit signing, verifies network access to VPN / GitHub / AWS / internal services, and ships a doctor command that verifies the setup and surfaces drift from baseline.
  • 2-hour engineering-platform walkthrough for your team, recorded — covers GitHub-as-code, the skills-repo workflow, and the onboarding CLI. Separate from the 90-minute IaC walkthrough in Core.

Governance and transparency baked in

  • Customer-managed KMS keys with explicit rotation control and a 30-day deletion window on every CMK. RDS, secrets, logs, and S3 default encryption all point at your CMKs, not AWS-managed keys.
  • S3 baseline — public-access block on, versioning on, lifecycle policy for noncurrent versions (90 days), Intelligent Tiering after 180 days. No embarrassing data leak, no surprise S3 bill.
  • Trivy ignore file is documented, not silent — every accepted IaC-scanner finding named with a rationale comment. No hidden security debt; what we accept, we name.
  • ECS Exec via SSM Session Manager — every ECS service ships with enableExecuteCommand wired and audit-logged. Zero SSH, zero bastion hosts, full session record.

Operational tooling

  • Incident-response runbook — security-incident triage template.
  • Drift-detection workflow, Terragrunt-managed, integrated with your CI/CD.
  • Monthly platform-health-report template.

Post-delivery support — 90 days

  • Drift review at day 30 and day 90 (each: 1 hour).
  • Slack channel for IaC questions — response within 1 business day, best-effort.
Read the full Core+ scope

P4

Container Platform — ECS

4–6 weeks fixed price, fixed scope

The deploy-one-application-end-to-end product. Three variants, one repeatable pattern.

Sold as one product because in practice nobody buys shared ECS infrastructure without an app to deploy on it. The substrate is built once per AWS account; the per-app surface is rebuilt on every additional deployment.

Three variants

New AWS account
5–6 weeks Substrate + first app.
Existing-substrate account
3–4 weeks Skip the substrate work if you already have VPC + cluster + shared ALB + KMS CMKs + Route53 at delivery quality.
Each additional app in the same account
2 weeks The substrate is reused, only the per-app surface is rebuilt.

What ships

ECS base infra — Route53, ALB, Certificates, Cluster

  • ECS Fargate cluster with capacity providers configured.
  • Shared Application Load Balancer — HTTPS listener with host-based routing. Multiple apps share the ALB — ~$25/mo amortized across N apps, not $25/mo per app.
  • Account-level Route53 hosted zone(s) for your primary domain, with A/AAAA aliases wired to CloudFront and to the ALB.
  • ACM certificates for the app domain, API domain, and marketing domain (us-east-1 for CloudFront; the ALB's region for the load balancer).
  • TLS 1.3 on the shared ALB — pinned to ELBSecurityPolicy-TLS13-1-2-Res-2021-06, with drop_invalid_header_fields = true and an HTTP → HTTPS 301 redirect baked in.
  • IAM baseline for ECS — task execution role base policy, service-linked roles, cluster-level permissions.
  • Account-level CloudWatch alarms — NAT bandwidth, ALB target health rollup, ALB 5xx rate aggregate.
  • Substrate runbooks — how to add a new app to the cluster, ALB host-routing patterns, ECR image lifecycle, IAM patterns for new-app onboarding.

Backend infra — ECR, ECS Services, AutoScaling, Secrets, Alarms

  • ECS Fargate service on the shared ALB (host-based routing rule for the app's hostname), HTTPS listener.
  • Autoscaling — CPU + memory target-tracking, min/max defaults set sensibly per your traffic expectation; configurable.
  • ECR repository for the app image, with lifecycle policy (untagged image cleanup, max image count) and image scanning enabled.
  • ECS migrations task — one-shot ECS task pattern runnable from CI before the service update. The same shape Comms and RuralOps run in production.
  • Authenticated container health check — the health endpoint sits behind a per-app API key on the internal port, not a public /health. Health-check observability lives outside the public attack surface.
  • Secrets Manager — per-service secret container, populated during the IaC handover.
  • Per-app KMS key for at-rest encryption of app-controlled resources; BYOK envelope-crypto pattern available on request.
  • SNS topic + optional PagerDuty webhook subscription — you provide the PagerDuty integration URL; we wire the HTTPS subscription.
  • ECS + ALB alarm pack — CPU high, memory high, no-healthy-tasks, 5xx rate, HealthyHostCount target health.
  • Three log-metric-filter patterns that catch what AWS-native alarms can't see: structured-JSON error level, plain-text fatal-at-startup, and partial-failure detection ("exit 0 but degraded" batch runs).
  • GitHub OIDC IAM role scoped at the ARN level, not policy-bag-of-stars. ecs:RunTask on the migrations task-definition ARN only; ecr:PutImage on the named app repo only; s3:* on the named frontend buckets only; cloudfront:CreateInvalidation on listed distribution IDs only; iam:PassRole only for the four named roles.
  • GitHub Actions reusable deploy workflow — PR build → main-push deploy → migrations → service update with rollback on health-check failure.

Frontend infra — CloudFront, S3 buckets

  • App frontend — CloudFront + S3 for the SPA, with SPA-friendly error mapping (403/404 → /index.html) and custom cache behaviors.
  • Public marketing-site frontend — second CloudFront + S3 distribution (on by default; removable at scoping if you don't need it).
  • CloudFront Origin Access Control (OAC), not deprecated OAI. TLSv1.2_2021 minimum. Managed CachingOptimized cache policy.
  • S3 buckets — application data (uploads, media) + frontend artifacts (app + public).

Async infra — Queues, async workers, scheduled jobs

  • Add-on. Included when workers or scheduled jobs are in scope; priced separately otherwise.
  • Background workers — SQS + Lambda resilience pattern: queue + DLQ, redrive after 5 receives, 14-day DLQ retention, visibility_timeout = lambda_timeout + 20s (no redelivery during processing), function_response_types = ["ReportBatchItemFailures"] (only failing records redrive).
  • Worker alarm pack — DLQ-has-messages alarm at threshold 0, Lambda errors / throttles / duration alarms.
  • Scheduled jobs (EventBridge → ECS RunTask) with alarm pack — not-firing, launch-errors, dropped, log-metric for job-level errors.

Data infra — RDS, RDS Proxy, DynamoDB, S3

  • Add-on. Included when the data tier is in scope; priced separately otherwise.
  • RDS (shared or dedicated) — per-workload security-group isolation, no subnet-wide ingress. The database accepts connections only from named security groups (service, migrations, scheduled-jobs, VPN).
  • Blue/Green RDS update path enabled — managed master password via Secrets Manager auto-rotation, Blue/Green-ready instance config.
  • IAM Database Authentication enabled — eliminates password rotation for DB auth.
  • RDS Proxy for the workloads that benefit (Lambda callers, connection-storm services).
  • DynamoDB tables — provisioned capacity or on-demand, per workload.

Fleet defaults + documentation

  • ARM64 Fargate by default — ~20% Fargate cost saving vs x86 on every task, no application code change required for compatible runtimes.
  • Capacity-provider mix — Fargate + Fargate Spot for non-critical workloads.
  • Architecture diagram for the app slice (and the substrate, on first-app-on-new-account engagements).
  • Runbooks per alarm — deploy, rollback, secret rotation, alert response, scale up/down, add a new env var, restart task, run a one-shot migration.

Optional add-ons

Twingate VPN connector, shared RDS, dedicated RDS, RDS Proxy, background workers (SQS + Lambda), scheduled jobs (EventBridge → ECS RunTask), DynamoDB tables, additional CloudFront + S3 frontends. Full pricing on the Container Platform — ECS page.

Read the full Container Platform — ECS scope

Some of the production applications we’ve built and currently operate.

Our own. Running on the same modules and the same live-config paths we ship to you. Real apps, real architectures, verifiable code.

  1. Comms

    A single HTTP API that sends Email and WhatsApp through one contract, one bill, and one place to debug delivery.

    Runs on the Core+ foundation and the Container Platform — ECS.

    Read the architecture
  2. ExpertCV

    Résumé and cover-letter generation grounded in a per-request EvidenceMap, so the LLM cannot fabricate.

    Runs on the Container Platform — ECS.

    Read the architecture
  3. PilatesTime

    B2B SaaS for pilates studios — PIX, multi-tenant RBAC, and a billing audit log that costs nothing extra to keep.

    Runs on the Core+ hardened foundation.

    Read the architecture
  4. RuralOps

    Bilingual ag-ops SaaS on Bun + Elysia, with multi-tenancy at the schema layer and AI assistant features that don't pretend to know things they don't.

    Runs on the Core foundation and the Container Platform — ECS.

    Read the architecture

Where is your AWS quietly costing you — in risk, ops, or dollars?

30 minutes with the founder. A 1-page PDF in your inbox within two business days. Eight dimensions, scored against the controls a production AWS account should clear.

We will be honest about whether you need us. If the gap is small enough that one of your engineers can close it in a Friday afternoon, we'll say so on the call.

Get the Scorecard

Free. No follow-up call unless you ask for one.

Skip the form. Talk to a person.

Plain email. No marketing reply, no scheduling tool, no funnel. Tell us what you're shipping and we'll tell you whether we're the right fit.