Well-Architected Framework
7 questions foundWhat is the AWS Well-Architected Framework and what are its six pillars?
Beginner The AWS Well-Architected Framework provides a consistent set of best practices for designing and evaluating cloud architectures, organized around six pillars including operational excellence, security, reliability, performance efficiency, cost optimization, and sustainability, each offering a set of guiding questions that help teams evaluate their workloads against proven architectural principles rather than relying purely on individual experience or guesswork.
// A Well-Architected review asks questions like:
// 'How do you back up data?' (Reliability)
// 'How do you monitor your resources to identify cost optimization opportunities?' (Cost Optimization)
Real-world example A development team building their first significant production application uses the Well-Architected Framework's guiding questions as a checklist to ensure they have not overlooked important considerations, such as proper backup strategies or least privilege security practices.
Common follow-ups: How often should a Well-Architected review be conducted for an existing workload?;Is the Well-Architected Framework specific to AWS or does it apply to cloud architecture generally?
AWS Backup & Disaster Recovery;IAM
What is the AWS Well-Architected Tool, and how does it help teams conduct a structured architecture review?
Beginner The AWS Well-Architected Tool is a free service available directly in the AWS console that guides you through a structured set of questions across all six pillars for a specific workload, automatically highlighting areas of potential risk based on your answers and providing specific improvement recommendations tied to relevant AWS documentation, making it easier to conduct a consistent, thorough architecture review without needing to manually track through the framework's guidance documents yourself.
aws wellarchitected create-workload --workload-name my-application --review-owner platform-team --lenses wellarchitected
Real-world example A platform team uses the Well-Architected Tool to review their newly launched application, and the tool immediately flags a high risk finding around their lack of automated backup testing, prompting them to address that gap before it becomes a real problem.
Common follow-ups: How does the Well-Architected Tool track improvement over multiple review cycles?;Are there specialized lenses available for specific industries or workload types?
AWS Backup & Disaster Recovery;AWS Config
How does the Operational Excellence pillar of the Well-Architected Framework guide teams in improving their processes and procedures?
Intermediate The Operational Excellence pillar focuses on running and monitoring systems to deliver business value while continuously improving processes and procedures, emphasizing practices such as performing operations as code through infrastructure automation, making frequent, small, reversible changes rather than large risky ones, anticipating failure through regular testing of failure scenarios, and learning from operational events and failures to continuously refine and improve processes over time.
// Operational excellence in practice
// Infrastructure defined as CloudFormation code
// Small, frequent deployments through automated CI/CD
// Blameless postmortems after every incident
Real-world example A team adopts a practice of conducting blameless postmortems after every production incident, systematically capturing lessons learned and concrete action items, steadily improving their operational maturity over time rather than repeating the same mistakes.
Common follow-ups: What does it mean to perform operations as code in practice?;How do you measure whether operational excellence is genuinely improving over time?
AWS CodePipeline CodeBuild & CodeDeploy (CI/CD);IaC (CloudFormation)
How does the Reliability pillar guide architectural decisions around fault tolerance, disaster recovery, and automatic recovery from failure?
Intermediate The Reliability pillar focuses on ensuring a workload performs its intended function correctly and consistently, guiding practices such as automatically recovering from failure by monitoring for signs of trouble and triggering automated responses, scaling horizontally to increase overall system availability, testing recovery procedures regularly rather than assuming they will work when actually needed, and managing change through automation to avoid the human error commonly introduced by manual configuration changes.
// Reliability principle in practice
// Auto Scaling Group automatically replaces unhealthy instances
// Multi-AZ RDS deployment for automatic database failover
Real-world example A company designs its application to automatically detect and replace unhealthy instances through an Auto Scaling Group, and separately conducts quarterly disaster recovery drills to validate that their database failover process actually works as expected under realistic conditions.
Common follow-ups: What is the difference between high availability and disaster recovery in the context of this pillar?;How do you calculate an appropriate recovery time objective for a specific workload?
AWS Backup & Disaster Recovery;Auto Scaling Groups
How does the Cost Optimization pillar guide teams toward avoiding unnecessary spending while still meeting business and performance requirements?
Intermediate The Cost Optimization pillar emphasizes practices such as adopting a consumption based pricing model where you pay only for what you actually use, measuring overall efficiency by relating business output to the associated cost, analyzing and attributing spending accurately to specific teams or workloads through consistent tagging, and continuously evaluating whether newer, more cost effective service options or architectural patterns have become available since the workload was originally designed.
aws ce get-cost-and-usage --time-period Start=2026-08-01,End=2026-09-01 --granularity MONTHLY --metrics UnblendedCost --group-by Type=TAG,Key=Project
Real-world example A team conducting a Well-Architected review discovers through the cost optimization pillar's guidance that they have been running a steady state workload entirely on On Demand pricing, and shifting a significant portion to Savings Plans immediately reduces their monthly bill.
Common follow-ups: How often should a workload be re evaluated for newer, more cost effective architectural options?;What is the relationship between the cost optimization pillar and the performance efficiency pillar?
AWS Cost Management & Billing;Tagging Strategies & Resource Management
How do you conduct an effective Well-Architected review for a complex, business critical workload, and what should happen with the findings afterward?
Advanced An effective review involves gathering the right stakeholders, including both technical architects and business owners who understand the workload's actual requirements, honestly answering each pillar's guiding questions rather than providing overly optimistic answers, prioritizing identified high risk items based on their potential business impact, creating a concrete remediation plan with clear ownership and timelines for each finding, and scheduling a follow up review to verify that improvements were actually implemented rather than letting the findings simply sit in a document that nobody revisits.
aws wellarchitected update-workload --workload-id abc123 --improvement-status IN_PROGRESS
Real-world example A company conducts a thorough Well-Architected review of its core payment processing workload, identifies several high risk findings related to disaster recovery testing, assigns clear ownership and a ninety day remediation timeline for each one, and schedules a follow up review to confirm the improvements were genuinely completed.
Common follow-ups: How do you prioritize which findings to address first when resources are limited?;What is a reasonable interval for conducting follow up reviews after an initial assessment?
AWS Backup & Disaster Recovery;AWS Cost Management & Billing
How can an organization use the Well-Architected Framework's specialized lenses to evaluate workloads against industry specific or technology specific best practices beyond the general six pillars?
Advanced Beyond the general Well-Architected Framework, AWS provides specialized lenses tailored to specific technology domains, such as the Serverless Lens, the Machine Learning Lens, and the SaaS Lens, along with lenses addressing specific industries, each adding additional guiding questions and best practices highly relevant to that particular context, letting an organization conduct a more deeply relevant review than the general framework alone would provide for a workload with very specific architectural characteristics.
aws wellarchitected create-workload --workload-name my-serverless-app --lenses wellarchitected serverless
Real-world example A company building an entirely serverless application applies the Serverless Lens alongside the standard Well-Architected Framework during their review, surfacing specific guidance around Lambda cold starts and event driven architecture patterns that the general framework alone would not have addressed as deeply.
Common follow-ups: How many specialized lenses does AWS currently offer, and how do you know which ones apply to your workload?;Can an organization create its own custom lens tailored to its specific internal standards?
Lambda & Serverless;Amazon SageMaker & Machine Learning on AWS