Overview of Talk Python to Me #559: 12 Things You Should (and Shouldn't) Do in AWS
In this episode, Michael Kennedy talks with Matt Lee of Schematical and Cloud War Games about practical AWS security and reliability habits that prevent late-night outages, expensive mistakes, and security incidents. The conversation focuses on a set of AWS do’s and don’ts for developers and teams, especially around infrastructure as code, IAM, networking, observability, deployment automation, cost control, and incident preparedness. They also discuss Cloud War Games, a training approach that simulates outages so teams can practice responding before a real 3 a.m. emergency hits.
Key Takeaways
- Most cloud disasters are preventable and usually trace back to convenience choices made earlier.
- The biggest AWS themes are:
- Use infrastructure as code
- Prefer roles over access keys
- Use least privilege everywhere
- Keep internal resources off public networks
- Automate deployments and logging
- Treat cost, security, and incident response as ongoing disciplines
- For smaller projects, you can still simplify a lot, but it’s wise to adopt the same habits early so they scale with you.
The 12 AWS Do’s and Don’ts
1. Don’t hand-provision infrastructure — use Terraform or CloudFormation
- Matt strongly recommends Infrastructure as Code (IaC).
- Tools mentioned:
- Terraform
- OpenTofu
- CloudFormation
- Benefits:
- Version control for infrastructure
- Repeatable environments
- Easy rebuilds after failures
- Less “mystery server” behavior
2. Don’t use access keys when you can use IAM roles
- IAM = Identity and Access Management
- Use roles attached to services like Lambda or EC2 instead of long-lived access keys.
- Access keys are risky because they can leak in
.envfiles, repos, or tooling. - Roles are more secure and easier to reason about in AWS.
3. Don’t use broad IAM permissions — use granular permissions
- Avoid
*/ wildcard permissions wherever possible. - Scope access to:
- specific actions
- specific buckets
- specific paths
- specific resources
- Principle: least privilege
- Their advice: don’t let convenience turn into “everything can do everything.”
4. Don’t put backend resources on public subnets
- Databases and internal services should generally live in private subnets.
- Public exposure should be reserved for things that must be internet-facing, such as:
- load balancers
- public APIs
- Matt emphasizes that private networking adds a critical extra layer of defense even if someone gets credentials.
5. Don’t use one security group for everything
- Security groups are AWS’s firewall-like controls.
- Each component should have tightly defined rules.
- Avoid “any-to-any” style rules that enable lateral movement if one system is compromised.
- Also consider restricting outbound traffic, not just inbound.
6. Don’t nurse EC2 instances — build immutable infrastructure
- “Cattle, not pets” was the theme here.
- Avoid hand-tuning servers after launch.
- Prefer:
- Docker
- Lambda
- ECS
- Benefits:
- repeatable deployments
- easy replacement
- better scaling
- less dependence on one special machine
7. Don’t rely on manual deployments — use CI/CD
- Matt recommends automated pipelines for:
- building images
- running tests
- running migrations
- deploying application changes
- AWS tools discussed:
- CodePipeline
- CodeBuild
- ECR (Elastic Container Registry)
- The point is to make deploys predictable and recoverable.
8. Don’t leave logging and observability as an afterthought
- Use CloudWatch logs, metrics, and alarms.
- Logging helps with:
- debugging outages
- understanding latency spikes
- tracking resource behavior
- Important caution:
- CloudWatch can become expensive if you query massive log ranges carelessly.
9. Don’t ignore cost controls
- Cost surprises are a major AWS pain point.
- Use:
- Cost Explorer
- budgets
- tagging
- Example discussed:
- hidden fees from outdated database versions
- runaway usage from poorly controlled services
- Core advice: monitor costs continuously, not just at month-end.
10. Don’t expose S3 buckets publicly
- Matt says public S3 buckets are a very common audit finding.
- Use:
- private buckets
- CloudFront signed URLs
- signed uploads when needed
- S3 is great for storage, but not as a public file server or CDN substitute at scale.
11. Don’t use EC2 bastions if you can use SSM Session Manager
- Bastion hosts are a traditional way to access private resources.
- Matt recommends SSM Session Manager instead when possible:
- no open SSH access
- no public bastion instance to maintain
- IAM-based access control
- If you do use a bastion, restrict it hard.
12. Don’t wait for the first incident to learn incident response
- Practice before the outage.
- Use simulations to make sure the team knows:
- who investigates
- how to communicate
- how to isolate impact
- how to recover quickly
Cloud War Games: Stress Inoculation for Outages
Matt’s Cloud War Games is a training concept built around simulated incidents and response drills.
What it is
- A controlled environment where outages, attacks, and failures are intentionally introduced.
- Designed to mimic the stress and ambiguity of real emergencies.
Why it matters
- Real incidents are chaotic.
- Newer engineers often freeze when pressure is high.
- Simulation helps teams practice:
- triage
- collaboration
- post-mortems
- escalation
- decision-making under stress
Additional uses
- Training internal teams
- Improving hiring evaluations
- Helping identify whether candidates:
- truly understand the stack
- can communicate under pressure
- work well with others
Additional Topics Discussed
AI and AWS safety
- They discussed the danger of giving AI agents direct access to production resources.
- Recommendation:
- use read-only tool calls
- route agent activity through safer layers
- prefer data lakes over production databases for heavy analysis
- Matt emphasized that agents are useful, but still need guardrails and human review.
Serverless and modern architecture choices
- For smaller or newer projects, Matt often suggests:
- Lambda + API Gateway
- private networking
- simple, focused AWS services
- His overall guidance: don’t try to use every AWS product; use a small, well-understood stack.
Kubernetes vs. Docker vs. ECS
- Matt is generally more comfortable recommending:
- Docker
- ECS
- He views Kubernetes as powerful but more complex, and often better suited to teams already committed to it.
Practical Recommendations
If you want the shortest possible AWS checklist from this episode:
- Use Terraform/OpenTofu/CloudFormation
- Use IAM roles, not long-lived access keys
- Apply least privilege everywhere
- Keep databases and internal systems private
- Use separate security groups per service
- Prefer Docker/Lambda/ECS over hand-managed servers
- Automate with CI/CD
- Turn on logging, metrics, and alerts
- Set budgets and cost alerts
- Keep S3 private
- Prefer SSM Session Manager over bastions
- Practice incidents with game days / war games
Closing Thought
The episode’s core message is simple: AWS problems are usually not solved at 3 a.m.; they’re prevented on ordinary afternoons by good defaults, strong automation, and disciplined security practices. Matt’s advice is aimed at making cloud systems easier to run, safer to operate, and less painful when something inevitably goes wrong.
