On this page
In short
- Use separate AWS accounts for production and non-production from day one, under AWS Organizations. Accounts are the strongest boundary AWS offers.
- People sign in through your identity provider with single sign-on. Nobody, human or pipeline, uses long-lived access keys.
- Every resource is defined in code (CDK, CloudFormation or Terraform) and deployed by a pipeline that authenticates to AWS with OIDC.
- Turn on audit logging, alarms and budget alerts before launch, not after the first surprise.
Why set up foundations for one app
It's tempting to create one AWS account, click together what the first app needs, and tidy up later. Later is easy to postpone. Production and test resources end up mixed together, nobody is sure which resources are safe to delete, the bill can't be split by app, and access has drifted to a few people with administrator rights.
A basic foundation isn't a large project. Most of it is configuration you set once: an account structure, sign-on, a pipeline pattern, logging and budgets. Done alongside the first app, it becomes the template every later app follows.
Accounts: the first decision
In AWS, an account is the hardest boundary between workloads: separate permissions, separate limits, separate bills. AWS Organizations groups accounts under one management account, with consolidated billing and policies that apply across them. A starting structure for a mid-sized company looks like this:
A starting account structure
Management
- Purpose
- Organizations, billing, single sign-on. No workloads.
- Who has access
- A very small number of administrators
Log archive / security
- Purpose
- Central audit logs and security tooling
- Who has access
- Security and a few administrators; read-mostly
Non-production
- Purpose
- Development and test environments for apps
- Who has access
- Developers, through the pipeline and SSO
Production
- Purpose
- Live apps only
- Who has access
- The pipeline; people read-only by default
| Account | Purpose | Who has access |
|---|---|---|
| Management | Organizations, billing, single sign-on. No workloads. | A very small number of administrators |
| Log archive / security | Central audit logs and security tooling | Security and a few administrators; read-mostly |
| Non-production | Development and test environments for apps | Developers, through the pipeline and SSO |
| Production | Live apps only | The pipeline; people read-only by default |
AWS Control Tower can set up this kind of structure with a baseline of controls applied for you; setting up Organizations directly is also reasonable at this size. Either way, add service control policies for a few firm rules, such as restricting which regions can be used and preventing anyone from turning off audit logging. As more apps arrive, you can split non-production and production further, per app or per team.
Identity and access
- Single sign-on for people. Connect AWS IAM Identity Center to your identity provider (Microsoft Entra ID, Okta or similar), so AWS access follows the same joiners-and-leavers process as everything else.
- Roles, not users. People get temporary credentials through permission sets per account: read-only, developer, administrator. No IAM users with access keys for day-to-day work.
- Lock down the root user of every account: strong password, MFA, and no access keys.
- Least privilege for apps. Each app's compute (Lambda functions, containers) gets its own IAM role with only the permissions it needs.
- Secrets in a secrets manager. Database passwords and third-party API keys live in AWS Secrets Manager or Parameter Store, never in code or environment files in the repository.
Infrastructure as code
Define every resource in code, review changes like any other code, and deploy them through a pipeline. That gives you a record of what exists and why, the ability to rebuild an environment, and identical dev and production setups. The three common choices:
Comparing common infrastructure-as-code tools on AWS
AWS CDK
- What it is
- Infrastructure in TypeScript, Python and other languages, deployed via CloudFormation
- Good fit when
- Developers own the infrastructure and want reusable components
CloudFormation
- What it is
- AWS's native declarative templates in YAML or JSON
- Good fit when
- You want plain templates with no extra tooling
Terraform
- What it is
- HashiCorp's declarative language with its own state file
- Good fit when
- You also manage other clouds or services, or already use it
| Tool | What it is | Good fit when |
|---|---|---|
| AWS CDK | Infrastructure in TypeScript, Python and other languages, deployed via CloudFormation | Developers own the infrastructure and want reusable components |
| CloudFormation | AWS's native declarative templates in YAML or JSON | You want plain templates with no extra tooling |
| Terraform | HashiCorp's declarative language with its own state file | You also manage other clouds or services, or already use it |
All three are sound. Pick one per company rather than per project, so everyone can read everyone else's infrastructure. Whichever you choose, the code should be able to create the whole environment from an empty account, rather than relying on resources someone once created by hand.
CI/CD with OIDC, not stored keys
A deployment pipeline needs permission to change AWS. The older approach stores an access key in the CI system, where it lives forever and can leak. The better approach is OpenID Connect (OIDC): GitHub Actions and most other CI systems can present a short-lived token that AWS trusts, and the pipeline assumes an IAM role for the length of the job.
Each account's deployment role trusts only the specific repository and branch or environment allowed to deploy there.
Pipeline checklist
- An OIDC identity provider in each workload account, and a deployment role whose trust policy names the exact repository and branch or environment.
- Separate roles for non-production and production; the production role can't be assumed from a feature branch.
- A required approval before production deploys, using your CI system's environment protection.
- Infrastructure changes shown as a diff in the pull request before they're applied.
- The same build artifact promoted from test to production, rather than rebuilt.
Environments
Most first apps need two environments, dev and production, defined by the same code with different settings. Add a staging environment when you need to test against production-like data or integrations before release. Keep environment differences in configuration, not in separate copies of the infrastructure code, so they can't quietly drift apart.
If the app has to reach systems in your own data center or plant, plan the connection early. A site-to-site VPN or AWS Direct Connect link, and the firewall rules on your side, usually involve people outside the project, so start those conversations early.
Monitoring and audit
- Audit trail. An organization-wide CloudTrail trail delivering to the log archive account, so management activity in every account (and any data events you choose to enable) is recorded somewhere app teams can't change.
- Threat detection. Amazon GuardDuty turned on across accounts, with findings routed to someone who will read them.
- App logs and metrics. CloudWatch logs with a set retention period per environment, and metrics for the things users feel: errors, latency, failed jobs.
- Alarms that reach people. CloudWatch alarms on those metrics, sent through SNS to email or your chat tool, with a named owner for each app.
- Backups. AWS Backup plans, or the database's own automated backups, with a restore tested at least once before launch.
Cost guardrails
Cloud costs surprise people when nobody is watching them, not because AWS is inherently expensive. A few settings make spending visible from the first day.
Cost guardrails
- AWS Budgets per account with alerts at thresholds finance agrees on, sent to both IT and the app owner.
- Cost Anomaly Detection turned on, to flag unusual spending patterns.
- Tags for app, environment and owner on every resource, applied by the infrastructure code and activated as cost allocation tags.
- Log retention and S3 lifecycle rules set, so storage doesn't grow forever.
- Serverless or scale-to-zero options for dev environments, and non-production resources shut down outside working hours where they can't scale to zero.
- A regular review of the bill by service and by tag, with the app owners involved.
Planning your first app on AWS?
Tell us what the app needs to do and what you already have in AWS, if anything. On a scoping call we'll talk through the foundation, the pipeline and what a fixed-price build would cover.
Not ready to talk yet? Draft a project brief first
Foundation checklist
Before the first app goes live
- Organizations set up with separate production and non-production accounts.
- IAM Identity Center connected to your identity provider; root users locked down.
- Infrastructure as code chosen and used for every resource.
- Pipeline deploying through OIDC roles, with an approval before production.
- CloudTrail, GuardDuty, alarms and backups in place and tested.
- Budgets, anomaly detection and cost allocation tags active.
- A named owner for each account and each app.
Our experience includes AWS infrastructure as code with CDK, CloudFormation and Terraform, CI/CD with GitHub Actions and OIDC, multi-environment deployments, and monitoring with CloudWatch. Round Table, our own product, runs on AWS with its infrastructure defined in AWS CDK, GitHub Actions for CI/CD and separate dev and production environments. See AWS cloud & DevOps for how we'd set up yours.
Sources
- What is AWS Organizations? and service control policies (AWS)
- What is AWS Control Tower? (AWS)
- What is IAM Identity Center? and root user best practices (AWS)
- What is AWS Secrets Manager? (AWS)
- AWS CDK developer guide (AWS)
- Configuring OpenID Connect in Amazon Web Services (GitHub)
- AWS Site-to-Site VPN (AWS)
- Creating a trail for an organization and GuardDuty with Organizations (AWS)
- What is AWS Backup? (AWS)
- AWS Budgets, Cost Anomaly Detection and activating cost allocation tags (AWS)
See it applied
- Illustrative sample. Not client work.See a sample scope for an internal tool →An invented request and approval portal that replaces a spreadsheet and email, on an existing SQL Server database with the company's existing single sign-on.
- Illustrative sample. Not client work.See a sample handover package →The handover index for an invented plant-floor dashboard: repositories, infrastructure as code, the pipeline, runbooks, a data dictionary, the credential handover list, alerts, known limitations and sign-off.
Share this guide
Related services
- AWS cloud & DevOps Infrastructure as code, deployment pipelines and monitoring for what we build.
- Custom software development The first app itself, built on a foundation the next ones can reuse.
- Web application development Dashboards, portals and internal tools built with React, Next.js and TypeScript.
Topics