<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Cloud Playbook]]></title><description><![CDATA[Playbooks for engineering leaders running multitenant SaaS on AWS who choose predictability over speed: boringly reliable services, controlled cloud spend, and audit-ready compliance.]]></description><link>https://www.thecloudplaybook.com</link><image><url>https://substackcdn.com/image/fetch/$s_!7MI5!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b1ca555-b578-4ae3-8dda-a03cbc6b1d18_500x500.png</url><title>The Cloud Playbook</title><link>https://www.thecloudplaybook.com</link></image><generator>Substack</generator><lastBuildDate>Wed, 29 Jul 2026 15:53:55 GMT</lastBuildDate><atom:link href="https://www.thecloudplaybook.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Amrut Patil]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[thecloudplaybook@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[thecloudplaybook@substack.com]]></itunes:email><itunes:name><![CDATA[Amrut Patil]]></itunes:name></itunes:owner><itunes:author><![CDATA[Amrut Patil]]></itunes:author><googleplay:owner><![CDATA[thecloudplaybook@substack.com]]></googleplay:owner><googleplay:email><![CDATA[thecloudplaybook@substack.com]]></googleplay:email><googleplay:author><![CDATA[Amrut Patil]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[TCP#136: You picked your secret store by cost. That was the wrong axis.]]></title><description><![CDATA[Secrets Manager vs. Parameter Store vs. KMS-encrypted env on AWS: a decision matrix. Why teams pick the wrong secret store by cost or habit, and the axes that produce a defensible choice per secret.]]></description><link>https://www.thecloudplaybook.com/p/secrets-manager-vs-parameter-store-aws</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/secrets-manager-vs-parameter-store-aws</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Sun, 26 Jul 2026 14:43:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Mwib!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14322c40-af7d-41ea-9db7-0893585ad472_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most teams choose a secret store the same way they choose a compute platform: by preference or cost, not by what the secret actually needs.</p><p>One team saw Secrets Manager&#8217;s per&#8209;secret pricing, decided it was expensive, and standardized on Parameter Store for everything. Another standardized on Secrets Manager because it sounded &#8220;purpose&#8209;built.&#8221; A third never decided at all:</p><ul><li><p>Database passwords in SSM parameters</p></li><li><p>API keys in Secrets Manager</p></li><li><p>A few credentials in KMS&#8209;encrypted env vars because that was fastest for one service two years ago</p></li></ul><p>Different routes, same destination: secrets live where habit or a cost comparison put them, not where the secret&#8217;s requirements put them.</p><p>And those requirements differ: a rotating database password is not the same as a static third&#8209;party API key, which is not the same as a config value that&#8217;s merely sensitive.</p><p>Optimizing for the cheapest monthly line item while ignoring rotation, audit, and incident response is optimizing the wrong variable.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Mwib!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14322c40-af7d-41ea-9db7-0893585ad472_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Mwib!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14322c40-af7d-41ea-9db7-0893585ad472_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!Mwib!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14322c40-af7d-41ea-9db7-0893585ad472_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!Mwib!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14322c40-af7d-41ea-9db7-0893585ad472_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!Mwib!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14322c40-af7d-41ea-9db7-0893585ad472_1448x1086.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Mwib!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14322c40-af7d-41ea-9db7-0893585ad472_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/14322c40-af7d-41ea-9db7-0893585ad472_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1559095,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/208409166?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14322c40-af7d-41ea-9db7-0893585ad472_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Mwib!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14322c40-af7d-41ea-9db7-0893585ad472_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!Mwib!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14322c40-af7d-41ea-9db7-0893585ad472_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!Mwib!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14322c40-af7d-41ea-9db7-0893585ad472_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!Mwib!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14322c40-af7d-41ea-9db7-0893585ad472_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h3>Why the Wrong Store Is a Real Cost</h3><p>A secret in the wrong store is not a billing issue. It is:</p><ul><li><p>A <span>rotation problem</span></p></li><li><p>An <span>audit problem</span></p></li><li><p>An <span>incident&#8209;response problem</span> waiting for its moment</p></li></ul><p><strong><span>Rotation cost:</span></strong><br>A database credential in plain Parameter Store does not rotate itself. Someone has to remember. They don&#8217;t. A credential that should live 30 days lives 2 years. When it leaks, the blast radius is 2 years of access, not 30 days.</p><p><strong><span>Audit cost:</span></strong><br>Compliance frameworks want proof: scheduled rotation, scoped access, logged reads. A secret in the wrong store cannot show that history. The auditor asks for rotation records, and there are none, because that store doesn&#8217;t rotate.</p><p><strong><span>Incident cost:</span></strong><br>When a credential leaks, the first questions are:</p><ul><li><p>How old is it?</p></li><li><p>What can it reach?</p></li><li><p>How fast can we rotate or revoke it?</p></li></ul><p>The store built for rotation answers quickly. The store chosen for price does not, and the incident stretches from hours into days.</p><p>The monthly price delta between stores is small.<br>The cost of a credential you cannot rotate, audit, or quickly revoke is very high, and shows up exactly when you can least afford it.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Cloud Playbook! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>How Teams End Up With Secrets Scattered</h3><p>Sprawl is not usually a decision. It&#8217;s the absence of one.</p><ul><li><p>The first secret goes &#8220;wherever&#8221; because one engineer had to ship one service.</p></li><li><p>The second secret goes somewhere else because another engineer made a different local call.</p></li><li><p>Neither choice is crazy in isolation. Together, they are the start of chaos.</p></li></ul><p>Then cost anxiety accelerates it:</p><ul><li><p>Secrets Manager&#8217;s per&#8209;secret line is visible on the bill</p></li><li><p>Parameter Store&#8217;s usage is fuzzier</p></li><li><p>A cost review flags Secrets Manager</p></li><li><p>The team moves secrets to the &#8220;cheaper&#8221; store without checking whether it meets the secrets&#8217; requirements</p></li></ul><p>The bill goes down. Rotation and audit capability go with it.</p><p>This persists because <span>nobody owns the decision</span>. There&#8217;s a standard for compute, a standard for networking, and no standard for secrets. So the question &#8220;where does this go?&#8221; gets answered per engineer, per service, per deadline.</p><p>Sprawl is the predictable result of a decision nobody was assigned to make.</p><div><hr></div><h3>The Axes That Actually Decide</h3><p>The &#8220;right&#8221; store is not about preference or price. It&#8217;s about four properties of the secret itself. Score the secret on these, and the store largely selects itself.</p><ol><li><p><span>Does it need automatic rotation?</span><br>This is the first cut.</p><ul><li><p>Must&#8209;rotate (e.g. most DB credentials, many service credentials) &#8594; Secrets Manager, with native rotation where possible</p></li><li><p>Does&#8209;not&#8209;rotate (static third&#8209;party key, certain config values) &#8594; does not need to pay for rotation it won&#8217;t use</p></li></ul></li><li><p><span>How sensitive is it and what&#8217;s the blast radius if it leaks?</span></p><ul><li><p>High&#8209;blast&#8209;radius (prod DB with customer data) &#8594; strongest access controls, logging, and rotation</p></li><li><p>Lower sensitivity (non&#8209;public but low&#8209;impact config) &#8594; can live in a simpler store</p></li></ul></li><li><p><span>How often is it read, and by what?</span><br>Read pattern affects both architecture and cost:</p><ul><li><p>Read on every request by a high&#8209;traffic service &#8594; needs caching and cost awareness</p></li><li><p>Read once at startup &#8594; throughput less relevant</p></li></ul><p>Parameter Store and Secrets Manager behave differently under load; the wrong choice here is either a latency issue or a surprise bill.</p></li><li><p><span>What&#8217;s its compliance scope?</span><br>In&#8209;scope for FedRAMP / HIPAA / ISO? Then you may be forced into:</p><ul><li><p>Demonstrable rotation</p></li><li><p>Scoped IAM</p></li><li><p>Auditable access history</p></li></ul><p>&#8220;We cannot demonstrate rotation&#8221; is a finding all by itself.</p></li></ol><p><span>KMS&#8209;encrypted environment variables</span> sit at the bottom of this matrix on purpose. They are right for a narrow case:</p><ul><li><p>Value needed only at process start</p></li><li><p>Not sensitive enough to justify a lookup per read</p></li><li><p>Does not rotate</p></li></ul><p>They are wrong for almost everything else. When secrets land here, it&#8217;s because no one applied the matrix at all.</p><div><hr></div><h3>What a Standard Changes</h3><p>Teams that define <span>where each class of secret belongs</span> stop putting secrets in the wrong place.</p><ul><li><p>Rotating database credentials live where rotation is native, so they actually rotate, and a leak&#8217;s blast radius is 30 days, not 2 years.</p></li><li><p>Static API keys live where they aren&#8217;t paying for rotation they don&#8217;t need.</p></li><li><p>High&#8209;sensitivity secrets live where access is tightly scoped and logged, so audits pull reports instead of doing archaeology.</p></li><li><p>Low&#8209;risk config values live in the simple store, so the expensive store is reserved for secrets that justify it.</p></li></ul><p>When a credential leaks, the response is fast because the store was chosen for that risk class. The team knows its age, reach, and rotation path because the store was designed to answer those questions.</p><p>Secrets stop being scattered by accident and start being placed by design. The store follows the secret&#8217;s requirements, and the requirements were explicitly examined instead of assumed.</p><div><hr></div><h2><strong>That&#8217;s it for today!</strong></h2><p>Did you enjoy this newsletter issue?</p><p>Share with your friends, colleagues, and your favorite social media platform.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Cloud Playbook</span></a></p><p><strong>Until next week &#8212; Amrut</strong></p><div><hr></div><h2><strong>Get in touch</strong></h2><p>You can find me on <a href="https://www.linkedin.com/in/patilamrut/">LinkedIn</a> or <a href="https://twitter.com/realamrutpatil">X</a>.</p><p>If you would like to request a topic to read, please feel free to contact me directly via LinkedIn or X.</p>]]></content:encoded></item><item><title><![CDATA[TCP #135: The GuardDuty org-wide gotchas that break your Terraform run]]></title><description><![CDATA[Delegated admin, per-feature enablement, severity routing, and the multi-region and provider traps, all as code]]></description><link>https://www.thecloudplaybook.com/p/guardduty-org-wide-terraform-aws</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/guardduty-org-wide-terraform-aws</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Thu, 23 Jul 2026 15:07:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!CrYV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6b5947-1c3a-4bc0-8ab1-007927a02c39_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You decided to enable GuardDuty org&#8209;wide as code. Good. Now you&#8217;re going to hit the gotchas that the console hides, but Terraform does not.</p><p>The console&#8217;s org&#8209;wide enablement is a smooth process because AWS paper-overs the underlying complexity. Terraform does not. When you enable GuardDuty across an organization with Terraform, you run into reality:</p><ul><li><p>GuardDuty is regional</p></li><li><p>Delegated administration has strict ordering</p></li><li><p>Auto&#8209;enable behaves differently for existing vs new accounts</p></li><li><p>Member association is not the same as feature enablement</p></li></ul><p>Teams that don&#8217;t know this up front get a rollout that half&#8209;works:</p><ul><li><p>GuardDuty on in one region, missing in three</p></li><li><p>New accounts get it, existing ones don&#8217;t (or the reverse)</p></li><li><p>Delegated administrator set, member accounts not associated</p></li></ul><p>Every one of these is a coverage gap that appears to be a success in the one dashboard you happen to be staring at.</p><p>This issue is the map of those gotchas and the code structure that avoids them.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CrYV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6b5947-1c3a-4bc0-8ab1-007927a02c39_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CrYV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6b5947-1c3a-4bc0-8ab1-007927a02c39_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!CrYV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6b5947-1c3a-4bc0-8ab1-007927a02c39_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!CrYV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6b5947-1c3a-4bc0-8ab1-007927a02c39_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!CrYV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6b5947-1c3a-4bc0-8ab1-007927a02c39_1448x1086.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CrYV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6b5947-1c3a-4bc0-8ab1-007927a02c39_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4c6b5947-1c3a-4bc0-8ab1-007927a02c39_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1473750,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/207499651?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6b5947-1c3a-4bc0-8ab1-007927a02c39_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CrYV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6b5947-1c3a-4bc0-8ab1-007927a02c39_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!CrYV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6b5947-1c3a-4bc0-8ab1-007927a02c39_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!CrYV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6b5947-1c3a-4bc0-8ab1-007927a02c39_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!CrYV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6b5947-1c3a-4bc0-8ab1-007927a02c39_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why the Gotchas Are Expensive</h2><p>A half&#8209;enabled detection system is more dangerous than an un&#8209;enabled one, because it creates false confidence.</p><p>The team believes GuardDuty is running org&#8209;wide. The audit evidence says it is enabled. But coverage has holes:</p><ul><li><p>A region never turned on</p></li><li><p>Accounts that predate the auto&#8209;enable setting</p></li><li><p>A feature enabled in the delegated administrator but never propagated</p></li></ul><p>The threat that lands in one of those holes is undetected, and nobody knows the hole exists until an incident review finds it. Discovering it after the fact means a multi&#8209;week backfill project instead of a plan and an apply.</p><p>The gotchas are cheap to handle before rollout and expensive after. Doing this as code only pays off if the code encodes the handling so the gaps never open.</p><h2>How the Gaps Open</h2><p>These coverage gaps are not due to carelessness. They are the predictable result of GuardDuty&#8217;s architecture meeting an org&#8209;wide rollout.</p><ul><li><p><strong><span>Regional service.</span></strong> GuardDuty is regional. Enabling it is a per&#8209;region action. Teams think in terms of &#8220;the organization,&#8221; focus on accounts, and then cover only their primary region while silently omitting the others.</p></li><li><p><strong><span>Auto&#8209;enable subtlety.</span></strong> Auto&#8209;enable controls whether <strong><span>new</span></strong> member accounts are enabled for GuardDuty. It does not retroactively turn it on for existing accounts. Teams that set auto&#8209;enable and assume full coverage miss every pre&#8209;existing account.</p></li><li><p><strong><span>Feature enablement vs association.</span></strong> Associating a member account with the delegated admin brings findings into view. It does not automatically enable additional protection features. Teams that associate all accounts and assume malware protection is on have coverage associated with accounts, not feature coverage.</p></li></ul><p>The console flow hides most of this. Terraform surfaces it. That&#8217;s a feature, not a bug: the code forces explicit decisions, and explicit decisions do not turn into silent gaps.</p><h2>The Decision Framework and the Gotchas</h2><p>Structuring Terraform correctly requires resolving four things. Each has a gotcha.</p>
      <p>
          <a href="https://www.thecloudplaybook.com/p/guardduty-org-wide-terraform-aws">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[TCP #134: GuardDuty org-wide is a one-click decision with a six-month tail.]]></title><description><![CDATA[The pre-flight decisions on delegated admin, cost, and alert routing that decide whether GuardDuty becomes signal or noise.]]></description><link>https://www.thecloudplaybook.com/p/guardduty-preflight-org-wide-aws</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/guardduty-preflight-org-wide-aws</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Sun, 19 Jul 2026 14:29:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!BWPk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4913c4d3-244a-48df-a9c4-3e02f25d049a_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong><span>Enabling GuardDuty across your organization takes just a few clicks. Making it useful is not.</span></strong></p><p>Most teams do the same thing:</p><ul><li><p>Someone in the security account enables GuardDuty</p></li><li><p>Delegated admin is set</p></li><li><p>Auto&#8209;enable is turned on for all member accounts</p></li></ul><p>Now the estate has threat detection &#8220;everywhere.&#8221; The dashboard fills with findings. Everyone feels safer.</p><p>Six months later, those findings land in an inbox nobody reads. The volume is high, the severity is mixed, and nobody has decided who acts on what.</p><p>GuardDuty is generating thousands of dollars <span>in value from&nbsp;</span><em><span>detection</span></em><span>&nbsp;and zero dollars in value from</span> <em>response</em>, because the response side was never designed.</p><p>GuardDuty is not the failure. GuardDuty works.<br>The failure is enabling a detection system without first deciding what happens when it detects something.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BWPk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4913c4d3-244a-48df-a9c4-3e02f25d049a_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BWPk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4913c4d3-244a-48df-a9c4-3e02f25d049a_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!BWPk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4913c4d3-244a-48df-a9c4-3e02f25d049a_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!BWPk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4913c4d3-244a-48df-a9c4-3e02f25d049a_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!BWPk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4913c4d3-244a-48df-a9c4-3e02f25d049a_1448x1086.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BWPk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4913c4d3-244a-48df-a9c4-3e02f25d049a_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4913c4d3-244a-48df-a9c4-3e02f25d049a_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1424922,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/207499125?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4913c4d3-244a-48df-a9c4-3e02f25d049a_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BWPk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4913c4d3-244a-48df-a9c4-3e02f25d049a_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!BWPk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4913c4d3-244a-48df-a9c4-3e02f25d049a_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!BWPk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4913c4d3-244a-48df-a9c4-3e02f25d049a_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!BWPk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4913c4d3-244a-48df-a9c4-3e02f25d049a_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Why the &#8220;Enable&#8221; Click Has a Long Tail</h2><p>GuardDuty is easy to turn on and expensive to turn on incorrectly. The expense is delayed enough that the connection is easy to miss.</p><p>You pay in two currencies:</p><p><strong><span>1. Dollars</span></strong></p><p>GuardDuty pricing scales with the volume of events it analyzes. Volume&#8209;heavy sources like:</p><ul><li><p>S3 data events</p></li><li><p>EKS audit logs</p></li><li><p>Malware protection</p></li><li><p>RDS protection</p></li></ul><p>Can multiply the bill in ways that surprise teams who enabled everything by default. An org&#8209;wide enablement with every feature on, across a large estate, can produce a monthly bill that triggers a finance conversation nobody planned for.</p><p><strong><span>2. Attention (the more damaging one)</span></strong></p><p>A detection system that produces more findings than the team can triage trains the team to ignore them.</p><p>The high&#8209;severity finding that matters arrives in the same flood as hundreds of low&#8209;severity informational findings, and it gets missed.</p><p>An ignored detection system is worse than no detection system because it creates the appearance of coverage without the substance.</p><p>Both costs are set by decisions made <em>before</em> the enable, not after. That makes this a pre&#8209;flight problem.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Cloud Playbook is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>How Good Teams Still Enable It Wrong</h2><p>The wrong&#8209;way enable is not negligence. It&#8217;s a reasonable response to how the feature presents itself.</p><p>GuardDuty&#8217;s setup flow makes the easy path the default path:</p><ul><li><p>Enable</p></li><li><p>Delegate</p></li><li><p>Auto&#8209;enable for all accounts</p></li><li><p>Done</p></li></ul><p>The flow does <strong><span>not</span></strong> ask:</p><ul><li><p>Who will read the findings?</p></li><li><p>What will this cost at your current event volume?</p></li><li><p>How does a critical finding reach a human with a pager?</p></li></ul><p>It just turns detection on, and that <em>feels</em> like the responsible thing to do.</p><p>Security teams are under pressure to show coverage.</p><p>&#8220;Is GuardDuty enabled?&#8221; is a question for an auditor and a customer.<br>The fastest way to answer &#8220;yes&#8221; is the one&#8209;click org&#8209;wide enable.</p><p>The box measures whether detection is on, not whether the response works.</p><p>So detection gets enabled comprehensively, and response is never designed. Findings accumulate. Volume becomes noise. The team that enabled a security control to reduce risk has added an alerting channel that everyone learns to mute.</p><p>The gap between &#8220;enabled&#8221; and &#8220;operational&#8221; is invisible until an incident falls into it.</p><div><hr></div><h2>Four Decisions to Make <em>Before</em> You Enable Org&#8209;Wide</h2><p>GuardDuty org&#8209;wide <em>is</em> worth doing. It&#8217;s worth doing <strong><span>after</span></strong> four decisions, not before.</p><p>Each one determines whether the system produces a signal or noise.</p><h3>1. Who is the delegated administrator?</h3><p>GuardDuty findings from every member account aggregate to a delegated administrator account.</p><p>That should be your <strong><span>security account</span></strong>, not the management account:</p><ul><li><p>Security tooling lives where the security team operates</p></li><li><p>The management account stays minimal and boring</p></li></ul><p>Choosing the wrong one is painful to reverse once findings are flowing.</p><h3>2. What will each feature cost at your scale?</h3><p>Foundational GuardDuty is one cost. The add&#8209;ons are separate:</p><ul><li><p>S3 protection</p></li><li><p>EKS audit log monitoring</p></li><li><p>Malware protection</p></li><li><p>RDS protection</p></li></ul><p>Each scale with your actual usage patterns.</p><p>Before enabling, estimate:</p><ul><li><p>Monthly cost <em>per feature</em></p></li><li><p>Against <em>your</em> current log and data volumes</p></li></ul><p>Then enable the features whose value justifies their cost. Not &#8220;all of them because they&#8217;re available.&#8221;</p><h3>3. Where do findings go, and who owns each severity?</h3><p>This is the decision that separates signal from noise.</p><p>Design this explicitly:</p><ul><li><p><strong><span>High</span></strong>: pages a human (on&#8209;call rotation/incident channel)</p></li><li><p><strong><span>Medium</span></strong>: goes to a queue where someone actually works (ticket system, triage channel)</p></li><li><p><strong><span>Low / Informational</span></strong>: logged for forensic use, not alerted</p></li></ul><p>Without this routing, every finding goes to the same destination with the same urgency. The important ones drown.</p><h3>4. What is your plan for the initial flood?</h3><p>The moment you enable GuardDuty org&#8209;wide, it surfaces findings for existing conditions. The first days produce a backlog, not a steady state.</p><p>Decide in advance:</p><ul><li><p>Who owns triage for the initial flood</p></li><li><p>How you&#8217;ll distinguish pre&#8209;existing conditions from new threats</p></li><li><p>Which classes of &#8220;old but still present&#8221; findings will trigger immediate work vs backlog remediation</p></li></ul><p>That keeps the launch from overwhelming the team on day one.</p><p><br>Enabling detection is a five&#8209;minute decision. Designing a response is the actual work.<br>Do the actual work first.</p><p>A detection system is only as valuable as the response it triggers.</p><div><hr></div><h2><strong>What Changes When Response Is Designed First</strong></h2><p>Teams that make these decisions <em>before</em> enabling get a security control that works, not a dashboard that decorates.</p><ul><li><p>Findings are routed by severity: critical ones go to a human, and informational ones go to a log.</p></li><li><p>The team trusts alerts because alerts are calibrated, and trusted alerts are acted on.</p></li><li><p>The monthly bill is a number someone chose, not a surprise that escalates.</p></li></ul><p>When an auditor asks about threat detection, the answer isn&#8217;t just &#8220;it&#8217;s enabled.&#8221;</p><p>The answer is: &#8220;It&#8217;s enabled, findings route to these owners at these severities, and here is our mean time to acknowledge a high&#8209;severity finding.&#8221;</p><p>That&#8217;s the difference between a checkbox and a capability.</p><p>GuardDuty stops being a system that generates findings nobody reads and becomes a system that generates responses that contain threats.</p><p>Detection was always the easy part.<br>Designing the response <em>before</em> you click enable is what makes the detection matter.</p><div><hr></div><h2>Coming Wednesday for paid subscribers</h2><p>On Wednesday, paid subscribers get the full system for enabling GuardDuty org&#8209;wide with Terraform, including the gotchas:</p><ul><li><p>Delegated administrator setup</p></li><li><p>Per&#8209;feature enablement decisions encoded as code</p></li><li><p>Severity&#8209;based routing and alerting design</p></li><li><p>Initial&#8209;flood triage plan</p></li><li><p>The Terraform and org&#8209;structure edge cases that break an org&#8209;wide enable</p></li></ul><p>If you want GuardDuty to be a control your team <em>runs</em>, not a dashboard you ignore, that&#8217;s the operating model.</p><div><hr></div><h3>Upgrade If You Need Implementation, Not Just Ideas</h3><p>If you&#8217;re using these emails to guide real decisions on your platform, you&#8217;ll get more leverage from the paid version of The Cloud Playbook.</p><p>The free newsletter gives you patterns and language.</p><p>The paid newsletter turns those patterns into implementation kits you can ship inside a quarter:</p><ul><li><p>Concrete rollout plans (90&#8209;day roadmaps for each pattern)</p></li><li><p>Templates and checklists (policies, runbooks, tagging schemes, review checklists)</p></li><li><p>Real examples from high&#8209;stakes AWS environments (what we actually shipped and why)</p></li></ul><p>If the paid side doesn&#8217;t save you more than the subscription in <strong>one</strong> incident, audit cycle, or bad migration you avoid, you should cancel and keep the playbooks.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe&quot;,&quot;text&quot;:&quot;Upgrade to the Paid Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thecloudplaybook.com/subscribe"><span>Upgrade to the Paid Cloud Playbook</span></a></p><div><hr></div><h2><strong>That&#8217;s it for today!</strong></h2><p>Did you enjoy this newsletter issue?</p><p>Share with your friends, colleagues, and your favorite social media platform.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Cloud Playbook</span></a></p><p><strong>Until next week &#8212; Amrut</strong></p><div><hr></div><h2><strong>Get in touch</strong></h2><p>You can find me on <a href="https://www.linkedin.com/in/patilamrut/">LinkedIn</a> or <a href="https://twitter.com/realamrutpatil">X</a>.</p><p>If you would like to request a topic to read, please feel free to contact me directly via LinkedIn or X.</p>]]></content:encoded></item><item><title><![CDATA[TCP #133: Permission sets belong in a pipeline, not in the Identity Center console]]></title><description><![CDATA[The controls-as-code system for permission sets: repo structure, CI/CD deployment, privilege review gates, and drift detection]]></description><link>https://www.thecloudplaybook.com/p/iam-identity-center-permission-sets-as-code</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/iam-identity-center-permission-sets-as-code</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Thu, 16 Jul 2026 15:02:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!rYsZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4e06a94-cc58-456f-8b53-00294b0afda8_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You centralized identity. IAM Identity Center is your source of truth. Human access runs through federation, not static keys. That was the hard part, and you did it.</p><p>Then someone widened a permission set in the console during an incident, and nobody knows.</p><p>The permission set, which was scoped to read-only, now has write access because an engineer needed it at 2 a.m., and the console was right there. The change is live in every account to which the permission set is assigned. There is no pull request, no reviewer, no record of why. The access review three months from now will find it, or it will not.</p><p>Centralizing identity solved the problem of credential sprawl. It did not solve the sprawl of permissions inside the centralized system.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rYsZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4e06a94-cc58-456f-8b53-00294b0afda8_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rYsZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4e06a94-cc58-456f-8b53-00294b0afda8_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!rYsZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4e06a94-cc58-456f-8b53-00294b0afda8_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!rYsZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4e06a94-cc58-456f-8b53-00294b0afda8_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!rYsZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4e06a94-cc58-456f-8b53-00294b0afda8_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rYsZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4e06a94-cc58-456f-8b53-00294b0afda8_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d4e06a94-cc58-456f-8b53-00294b0afda8_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1478943,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/206649627?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4e06a94-cc58-456f-8b53-00294b0afda8_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rYsZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4e06a94-cc58-456f-8b53-00294b0afda8_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!rYsZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4e06a94-cc58-456f-8b53-00294b0afda8_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!rYsZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4e06a94-cc58-456f-8b53-00294b0afda8_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!rYsZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4e06a94-cc58-456f-8b53-00294b0afda8_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why Console-Edited Permission Sets Undo the Work</h2><p>The value of IAM Identity Center lies in the fact that access is defined in one place. That value evaporates the moment the one place is edited by hand.</p><p>A permission set is not a small object. It is a bundle of policies that applies to potentially dozens of accounts at once. A single console edit to a widely assigned permission set is one of the highest-blast-radius changes available in your entire AWS estate, and it is available to anyone with console access to the management account.</p><p>The cost is the same as that of any unversioned control. At audit time, you cannot show the history of who changed what and why. During an incident, you cannot tell whether a permission set is in its intended state or in the state someone left it months ago. And the privilege-escalation review that your compliance framework requires becomes impossible, because the thing you are reviewing changes underneath you without a trace.</p><p>You did the work to centralize. Managing permission sets by hand quietly gives it back.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Cloud Playbook is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>How Teams Backslide After Centralizing</h2><p>The backslide is not a decision. It is the path of least resistance reasserting itself.</p>
      <p>
          <a href="https://www.thecloudplaybook.com/p/iam-identity-center-permission-sets-as-code">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[TCP #132: Your Control Tower guardrails belong in Terraform, not the console]]></title><description><![CDATA[The controls-as-code system for preventive, detective, and proactive guardrails, with review, drift detection, and rollout order.]]></description><link>https://www.thecloudplaybook.com/p/control-tower-controls-as-code-terraform</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/control-tower-controls-as-code-terraform</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Sun, 12 Jul 2026 14:25:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!PhQW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5e520ec-d597-4163-afdc-1fb472df984f_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most teams enable Control Tower controls by clicking through the console.</p><p>Someone enables a guardrail during setup. Someone else enables three more during an audit scramble. A year later, the landing zone has 40 controls applied across five OUs, and no one can tell you why any single control is on, who turned it on, or what breaks if it comes off.</p><p>The controls work. They deny the actions they are supposed to deny. But the configuration lives in the console, not in a repository, and that means it cannot be reviewed, diffed, rolled back, or reasoned about as a system.</p><p>This is the gap between having guardrails and operating them.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PhQW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5e520ec-d597-4163-afdc-1fb472df984f_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PhQW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5e520ec-d597-4163-afdc-1fb472df984f_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!PhQW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5e520ec-d597-4163-afdc-1fb472df984f_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!PhQW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5e520ec-d597-4163-afdc-1fb472df984f_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!PhQW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5e520ec-d597-4163-afdc-1fb472df984f_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PhQW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5e520ec-d597-4163-afdc-1fb472df984f_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e5e520ec-d597-4163-afdc-1fb472df984f_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1516912,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/206648884?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5e520ec-d597-4163-afdc-1fb472df984f_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PhQW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5e520ec-d597-4163-afdc-1fb472df984f_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!PhQW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5e520ec-d597-4163-afdc-1fb472df984f_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!PhQW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5e520ec-d597-4163-afdc-1fb472df984f_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!PhQW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5e520ec-d597-4163-afdc-1fb472df984f_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why Console-Managed Controls Become a Liability</h2><p>Control Tower controls are the enforcement layer of your landing zone. They are also, for most teams, the least version-controlled part of the entire AWS estate.</p><p>The cost shows up at audit time. An auditor asks which controls enforce a given SOC 2 requirement, and the team produces a screenshot instead of a commit. The evidence is a point-in-time snapshot with no history, no author, and no justification.</p><p>The cost also shows up during incidents. A deploy fails because a proactive control blocks it. The engineer does not know the control exists, cannot find where it is configured, and cannot tell whether disabling it is safe. The control meant to prevent a mistake now causes an outage in developer velocity instead.</p><p>Controls managed by clicking are controls that drift. Someone disables one to unblock a launch and never re-enables it. The guardrail that your compliance posture depends on is now off, and nothing in your system knows.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Cloud Playbook is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>How Smart Teams End Up Here</h2><p>No one decides to manage controls by hand. The console path is simply the fastest way to unblock the immediate need.</p><p>Control Tower&#8217;s own interface encourages it. The landing zone is set up through the console, so the first controls are enabled there. The pattern is set before anyone asks whether it should be code.</p><p>Then the team scales. More OUs, more accounts, more frameworks in scope. Each new requirement adds controls, and each addition happens where the last one did: in the console, under deadline, by whoever is closest to the problem.</p><p>The team that would never deploy a Lambda function without a repository is managing the enforcement layer of its entire compliance posture with mouse clicks. The inconsistency is invisible because controls feel like configuration rather than infrastructure. They are infrastructure. They deny real actions in real accounts, and they belong under the same discipline as everything else you deploy.</p><h2>The Three Control Types and How to Manage Each</h2><p>Control Tower exposes three control types, and the controls-as-code decision varies by type.</p><ol><li><p><strong>Preventive controls</strong> are Service Control Policies. They stop an action before it happens: deny disabling CloudTrail, deny resource creation outside approved regions. These are the highest-value controls to manage as code because they are the ones an incident is most likely to touch. Manage them in Terraform against the OU, with the policy body in the repository so the exact denied actions are reviewable.</p></li><li><p><strong>Detective controls</strong> are AWS Config rules. They evaluate resources after creation and flag non-compliance, such as an S3 bucket without encryption or a security group with open ingress. Manage these as code so the rule set is consistent across all accounts, and so that a new framework requirement becomes a pull request, not a console session.</p></li><li><p><strong>Proactive controls</strong> are CloudFormation Guard hooks. They evaluate a resource before provisioning and block deployment if the evaluation fails. These are the controls most likely to surprise a developer, so managing them as code matters most for a different reason: the repository is where an engineer looks to understand why their deployment was blocked.</p></li></ol><p>The axis to weigh is the blast radius against reversibility. A preventive control at the org root has the widest blast radius and the slowest reversal. It gets the most reviews. A detective control in a single sandbox OU has a narrow blast radius and instant reversal. It gets less. Match the review weight to the control&#8217;s reach, and do not apply the same ceremony to every control regardless of consequence.</p><p>The second axis is framework mapping. Every control should trace to a named requirement. A control that enforces nothing anyone can name is a control that will eventually break a deploy for a reason no one can defend.</p><h2>Implementation Steps</h2><p>Deploy controls as code in this order. The sequence matters because early mistakes are the expensive ones.</p><ol><li><p><strong>Inventory the current state first.</strong> Before writing any Terraform, export every control currently enabled, the OU it targets, and its type. This is your baseline. You cannot manage as code what you have not first written down.</p></li><li><p><strong>Map each control to a requirement.</strong> For every control in the inventory, name the framework requirement or internal policy it enforces. Controls that map to nothing are candidates for removal, not import.</p></li><li><p><strong>Import, do not recreate.</strong> Bring existing controls into Terraform state using import blocks rather than destroying and recreating them. Recreating a preventive control means there is a window when the guardrail is off. Import avoids the gap.</p></li><li><p><strong>Structure the repository by OU and type.</strong> One module per OU, controls grouped by preventive, detective, and proactive. The structure should allow a reviewer to see all controls for a given OU in one place.</p></li><li><p><strong>Enforce review by blast radius.</strong> Controls at the org root or a production OU require a two-person review. Controls on a sandbox OU can ship with one. Encode this in the pull request rules, not in a wiki no one reads.</p></li><li><p><strong>Add drift detection.</strong> Run a scheduled plan against the control configuration. Any drift between the repository and the deployed state pages the owner. This is the check that catches the disabled-and-forgotten control before an auditor does.</p></li><li><p><strong>Stage rollout through OU progression.</strong> A new control lands on the sandbox OU first, then a non-production OU, then production. Never apply a new preventive control to the org root as its first deployment.</p></li></ol><h2>What to Measure</h2><p>Controls as code is a system, and the system should be measured.</p><ol><li><p><strong>Control coverage.</strong> The percentage of enabled controls that exist in the repository versus the console. Track this weekly during migration. The target is 100 percent, and anything managed in the console after migration is a regression.</p></li><li><p><strong>Drift events per month.</strong> The count of times the deployed control state diverged from the repository. A healthy landing zone trends toward zero. Persistent drift signals that someone is still editing controls by hand.</p></li><li><p><strong>Time-to-evidence for a control question.</strong> How long does it take to answer &#8220;which controls enforce this requirement&#8221; during an audit? With controls as code, this repository query takes minutes to run. An above-an-hour means the framework mapping is incomplete.</p></li></ol><p>Review these in the platform's monthly operating review. Coverage and drift are the two metrics that indicate whether the enforcement layer is under control.</p><h2>What Changes When Controls Are Code</h2><p>The landing zone stops being a black box.</p><p>Every guardrail has an author, a justification, and a history. The audit question that used to take a day of screenshotting becomes a repository search. The control that blocks a deploy is documented where the engineer already looks. The disabled-and-forgotten guardrail is caught by drift detection within a day, rather than surfacing at the next audit.</p><p>The enforcement layer of your compliance posture becomes a system you operate, not a configuration you hope is still correct.</p><div><hr></div><h3><strong>Upgrade If You Need Implementation, Not Just Ideas</strong></h3><p>If you&#8217;re using these emails to guide real decisions on your platform, you&#8217;ll get more leverage from the paid version of The Cloud Playbook.</p><p>The free newsletter gives you patterns and language.</p><p>The paid newsletter turns those patterns into implementation kits you can ship inside a quarter:</p><ul><li><p>Concrete rollout plans (90&#8209;day roadmaps for each pattern)</p></li><li><p>Templates and checklists (policies, runbooks, tagging schemes, review checklists)</p></li><li><p>Real examples from high&#8209;stakes AWS environments (what we actually shipped and why)</p></li></ul><p>If the paid side doesn&#8217;t save you more than the subscription in <strong>one</strong> incident, audit cycle, or bad migration you avoid, you should cancel and keep the playbooks.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe&quot;,&quot;text&quot;:&quot;Upgrade to the Paid Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thecloudplaybook.com/subscribe"><span>Upgrade to the Paid Cloud Playbook</span></a></p><div><hr></div><h2><strong>That&#8217;s it for today!</strong></h2><p>Did you enjoy this newsletter issue?</p><p>Share with your friends, colleagues, and your favorite social media platform.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Cloud Playbook</span></a></p><p><strong>Until next week &#8212; Amrut</strong></p><div><hr></div><h2><strong>Get in touch</strong></h2><p>You can find me on <a href="https://www.linkedin.com/in/patilamrut/">LinkedIn</a> or <a href="https://twitter.com/realamrutpatil">X</a>.</p><p>If you would like to request a reading topic, please feel free to contact me directly via LinkedIn or X.</p>]]></content:encoded></item><item><title><![CDATA[TCP #131: The multi-account checklist regulated SaaS teams should have written]]></title><description><![CDATA[Account strategy, identity, networking, tenancy, and audit &#8212; each section with owners, cadences, and evidence.]]></description><link>https://www.thecloudplaybook.com/p/aws-multi-account-design-checklist-saas</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/aws-multi-account-design-checklist-saas</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Thu, 09 Jul 2026 15:06:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!oiJ-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff736cf3-9dff-4458-903c-0ce21a601019_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>AWS publishes excellent reference architectures for multi-account design. Most engineering teams have read them. Most engineering teams still ship multi-account environments that fail audits, leak access, or sprawl into operational debt.</p><p>The reference architectures are not wrong. They are incomplete in a specific way: they describe what the design looks like, not what the team commits to in order for the design to stay intact 18 months from now.</p><p>A checklist closes that gap. It explicitly names each design decision, the team that owns it, the verification cadence that prevents drift, and the artifact that proves the design is still in effect. The checklist does not produce a better design than the reference architecture. It produces a design that survives the conditions that erode reference architectures: org-chart changes, scope expansion, deadline pressure, and the silent drift that comes from boundaries without owners.</p><p>The checklist below is the one I would use with a regulated SaaS platform team. </p><p>It assumes SOC 2 Type II at a minimum, with a roadmap toward FedRAMP Moderate, HIPAA, or ISO 27001. The structure applies even if the team operates outside those frameworks; the cadences and evidence requirements are tighter for regulated workloads.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oiJ-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff736cf3-9dff-4458-903c-0ce21a601019_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oiJ-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff736cf3-9dff-4458-903c-0ce21a601019_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!oiJ-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff736cf3-9dff-4458-903c-0ce21a601019_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!oiJ-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff736cf3-9dff-4458-903c-0ce21a601019_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!oiJ-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff736cf3-9dff-4458-903c-0ce21a601019_1448x1086.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oiJ-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff736cf3-9dff-4458-903c-0ce21a601019_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ff736cf3-9dff-4458-903c-0ce21a601019_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1991753,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/205060366?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff736cf3-9dff-4458-903c-0ce21a601019_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oiJ-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff736cf3-9dff-4458-903c-0ce21a601019_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!oiJ-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff736cf3-9dff-4458-903c-0ce21a601019_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!oiJ-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff736cf3-9dff-4458-903c-0ce21a601019_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!oiJ-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff736cf3-9dff-4458-903c-0ce21a601019_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>How to Use the Checklist</h2><p>The checklist has seven sections. Each section contains design items, owners, verification cadences, and required artifacts.</p><p>For each item, the team commits to four things:</p>
      <p>
          <a href="https://www.thecloudplaybook.com/p/aws-multi-account-design-checklist-saas">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[TCP #130: Multi-Account AWS Environments Fail When Nobody Owns the Boundaries]]></title><description><![CDATA[Why account sprawl, tenant separation, and shared services break down without explicit control-plane ownership.]]></description><link>https://www.thecloudplaybook.com/p/multi-account-aws-boundaries-ownership</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/multi-account-aws-boundaries-ownership</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Sun, 05 Jul 2026 14:28:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!WhFP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1f9ead6-108d-4126-9eea-ebae82bcf125_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most engineering organizations adopted multi-account AWS for the right reasons.</p><p>The single-account deployment was sprawling. The blast radius was unacceptable. Compliance was difficult to scope. </p><p>The team split workloads across accounts: production, staging, and dev. Then a separate account for the data team. Then accounts for individual customer tenants. Then a logging account. Then a security tooling account.</p><p>By year three, the organization has 40 accounts. AWS Organizations holds the top of the tree. Control Tower, if the team adopts it, holds the bottom. The middle is unclear.</p><p>The accounts work. Workloads run in them. Bills are paid. But the boundaries between accounts, the rules about what crosses what and who decides, are owned by nobody.</p><p>When something goes wrong, the cost is visible. The data engineer cannot get cross-account access to the bucket they need. The compliance team cannot find the audit log for an action in the production account. The platform team discovers that three accounts have public S3 buckets nobody owns. The security team flags an IAM role assumption pattern that has been running for 14 months without review.</p><p>The accounts are not the problem. The boundaries between them are. And boundaries without owners produce the same outcomes as no boundaries at all.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WhFP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1f9ead6-108d-4126-9eea-ebae82bcf125_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WhFP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1f9ead6-108d-4126-9eea-ebae82bcf125_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!WhFP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1f9ead6-108d-4126-9eea-ebae82bcf125_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!WhFP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1f9ead6-108d-4126-9eea-ebae82bcf125_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!WhFP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1f9ead6-108d-4126-9eea-ebae82bcf125_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WhFP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1f9ead6-108d-4126-9eea-ebae82bcf125_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d1f9ead6-108d-4126-9eea-ebae82bcf125_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1502056,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/205055793?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1f9ead6-108d-4126-9eea-ebae82bcf125_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!WhFP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1f9ead6-108d-4126-9eea-ebae82bcf125_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!WhFP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1f9ead6-108d-4126-9eea-ebae82bcf125_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!WhFP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1f9ead6-108d-4126-9eea-ebae82bcf125_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!WhFP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1f9ead6-108d-4126-9eea-ebae82bcf125_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Why Account Sprawl Happens</h2><p>Multi-account architectures rarely start sprawling. They become sprawling because account creation is rewarded and account boundaries are not.</p><p>The team that needs isolation creates an account. Provisioning is fast; the security team approves it, and the workload comes in. The team is unblocked. Nobody asks who owns the account in 18 months, when the workload changes shape, who decides which other accounts can call into it, who reviews the IAM policies written into it, or who confirms that logging is exported to the central account.</p><p>The account becomes orphaned at the moment it is provisioned. The team that owns the workload owns the workload. The platform team owns the foundation. Nobody owns the account.</p><p>Multiply this by 18 months of provisioning decisions. The org chart shifts twice. The platform team turns over a third of its engineers. The compliance scope expands to a new framework. The original team that requested the account no longer remembers why they needed it.</p><p>The boundary that was clear at creation time is now ambiguous. The account is still running. The workload inside it is operationally healthy. The boundary between this account and the others is undefined, and the security team has no record of who would answer questions about it.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Cloud Playbook is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>What Boundary Ownership Actually Means</h2><p>Boundary ownership is not the same as account ownership.</p><p>The team that operates a workload owns the workload. They are responsible for the application, deployment, on-call rotation, and cost. This is straightforward, and most organizations get it right.</p><p>Boundary ownership is the layer above. It governs the rules about what crosses the account, in either direction. Specifically:</p><ul><li><p><strong>Identity boundary.</strong> Who can assume IAM roles in this account from outside, and what those roles can do? Who can assume IAM roles into other accounts from this one, and what those roles can do. Trust relationships, condition keys, session policies.</p></li><li><p><strong>Network boundary.</strong> What VPC peering, Transit Gateway attachments, or PrivateLink endpoints exist into and out of this account? What CIDR ranges are routable? What ports are open at the security group level for cross-account traffic?</p></li><li><p><strong>Data boundary.</strong> What S3 buckets, KMS keys, or other resources in this account can be accessed from other accounts? What buckets, keys, or resources in other accounts can this account access? Resource policies, bucket policies, KMS key policies.</p></li><li><p><strong>Audit boundary.</strong> What CloudTrail, Config, GuardDuty, and Security Hub data flows out of this account into the central audit account? Whether the flow is enforced, whether it is verified, and who reviews it.</p></li><li><p><strong>Cost boundary.</strong> Which cost center is the spend in this account allocated to? Whether the allocation is automatic via tagging policy or manual via accounting reconciliation. Who reviews variance?</p></li></ul><p>Each of these boundaries has a different ownership profile. The team that runs the workload does not own the identity boundary. The platform team does. The compliance team does not own the network boundary. The networking team does. When ownership is explicitly named for each boundary, the team has a forwarding address for boundary-related questions. When ownership is implicit, the questions sit in nobody&#8217;s queue.</p><div><hr></div><h2>The Three Failure Patterns of Boundary Drift</h2><p>Boundary ownership ambiguity produces three predictable failure patterns.</p><ol><li><p><strong>The shadow trust relationship.</strong> A workload in account A needs to read a bucket in account B. The team writes a bucket policy. The team writes a role with the right permissions. The role works. Nobody flags that the bucket policy now grants cross-account access. Six months later, a security review discovers the policy and cannot trace why it was written or whether it is still needed. Removing it might break the workload. Leaving it leaves an undocumented trust path. The team chooses to leave it. The trust path stays.</p></li><li><p><strong>The audit gap.</strong> A new account is provisioned. The team configures the workload. The team forgets to enable CloudTrail forwarding to the central audit account. Eleven months later, the audit team is collecting evidence for SOC 2 Type II and discovers that the account's CloudTrail data has been missing for 11 months. The audit finding is significant. The remediation is a one-line configuration change that should have been made on day zero. Nobody owned the audit boundary, so nobody verified it.</p></li><li><p><strong>The networking debt.</strong> Each new account is given a Transit Gateway attachment. The attachments accumulate. The CIDR ranges overlap or nearly overlap. The platform team does not realize this until a new account cannot peer because the CIDR space is exhausted. The fix is a re-IP project that takes a quarter. The cause was a boundary nobody was reviewing as accounts were created.</p></li></ol><p>These patterns are not specific to one organization. They are predictable outcomes of the same root cause: boundaries without explicit owners drift, and drift compounds.</p><div><hr></div><h2>Four Questions That Surface Whether Your Boundaries Have Owners</h2><p>Before the next account is provisioned, the platform team should be able to answer four questions in writing. If they cannot, boundary ownership is the work that needs to come before more accounts.</p><ol><li><p><strong>For each existing account, who owns the identity boundary?</strong> The named individual or team that approves cross-account role assumptions, reviews trust policies, and decides when a role can be assumed from a new principal. If the answer is &#8220;the platform team&#8221; without a specific person and a documented review cadence, the answer is no.</p></li><li><p><strong>For each existing account, who owns the network boundary?</strong> The named individual or team that approves VPC peering, Transit Gateway attachments, and PrivateLink endpoints. The team that maintains the CIDR allocation registry. If the team cannot produce the registry, the boundary is unowned.</p></li><li><p><strong>For each existing account, who verifies that audit data is flowing to the central account?</strong> Not configured. Verified. The difference is monthly: configuration drifts, forwarding breaks, and retention policies change. The team that owns audit boundary verification runs a monthly check. If no team runs the check, the audit boundary is unowned.</p></li><li><p><strong>When a new account is provisioned, what is the first action that establishes ownership of each boundary?</strong> A defined process that names the identity boundary owner, the network boundary owner, the audit boundary verification owner, and the cost boundary owner before the account is handed to the workload team. If account provisioning does not include this step, every new account starts unowned.</p></li></ol><p>These four questions are diagnostic. The honest answer reveals which boundaries are owned, which are partially owned, and which are not owned at all. Most organizations discover that one or two of the four are well known, while the others are not.</p><div><hr></div><h2>What Changes When Boundaries Have Owners</h2><p>When boundary ownership is explicitly named, the failure patterns above stop occurring.</p><p>The shadow trust relationship is caught at creation. The identity boundary owner reviews the trust policy as part of the workload onboarding. Cross-account access has a documented justification, a documented review cadence, and a removal trigger if the workload changes.</p><p>The audit gap is caught within 30 days. The audit boundary owner runs a monthly verification. Configuration drift is detected before it becomes an audit finding. The CloudTrail forwarding question is answered before the auditor asks.</p><p>The networking debt is caught at the design phase. The network boundary owner maintains the CIDR registry and reviews each new attachment against the existing topology. Re-IP projects do not become quarter-long initiatives because the topology was managed instead of accumulated.</p><p>The platform team&#8217;s relationship to multi-account architecture changes. The team is no longer reactive to drift. The team is proactive about the boundary that prevents drift. The accounts continue to multiply as the workload portfolio grows. The boundaries stay coherent because they are owned.</p><div><hr></div><h2>Why This Is the Right Conversation Now</h2><p>The teams most exposed to boundary drift are SaaS organizations expanding into regulated customer segments.</p><p>A SOC 2 Type II audit produces findings on boundary drift. A FedRAMP authorization fails on it. A HIPAA business associate agreement requires the organization to demonstrate control over the boundary, not just over the workload. A customer requesting tenant isolation in a separate account requires the organization to have a defensible model for what crosses the account boundary and what does not.</p><p>The cost of unowned boundaries is invisible until one of these events. By that point, the cost is high.</p><p>Teams that name boundary ownership before the next account is provisioned do not incur these costs in the same way. The boundary is owned. The boundary is reviewed. The boundary survives org-chart changes because it is documented. The audit, the customer, and the framework all encounter a defended boundary, not a boundary that has been drifting for two years.</p><div><hr></div><p>On Wednesday, <strong>paid subscribers get the AWS Multi-Account Design Checklist for regulated SaaS platforms</strong>. The checklist covers account strategy, logging, identity, networking, tenant isolation, shared services, and audit evidence. Each section names the explicit owner, the verification cadence, and the artifact that confirms the boundary is intact. Use it to baseline your existing accounts before the next audit, or as the design document before the next account is provisioned.</p><p><strong>Upgrade If You Need Implementation, Not Just Ideas</strong></p><p>If you&#8217;re using these emails to guide real decisions on your platform, you&#8217;ll get more leverage from the paid version of The Cloud Playbook.</p><p>The free newsletter gives you patterns and language.</p><p>The paid newsletter turns those patterns into implementation kits you can ship inside a quarter:</p><ul><li><p>Concrete rollout plans (90&#8209;day roadmaps for each pattern)</p></li><li><p>Templates and checklists (policies, runbooks, tagging schemes, review checklists)</p></li><li><p>Real examples from high&#8209;stakes AWS environments (what we actually shipped and why)</p></li></ul><p>If the paid side doesn&#8217;t save you more than the subscription <span>cost in&nbsp;</span><strong><span>a single</span></strong> incident, audit cycle, or bad migration you avoid, you should cancel and keep the playbooks.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe&quot;,&quot;text&quot;:&quot;Upgrade to the Paid Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thecloudplaybook.com/subscribe"><span>Upgrade to the Paid Cloud Playbook</span></a></p><div><hr></div><h2><strong>That&#8217;s it for today!</strong></h2><p>Did you enjoy this newsletter issue?</p><p>Share with your friends, colleagues, and your favorite social media platform.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Cloud Playbook</span></a></p><p><strong>Until next week &#8212; Amrut</strong></p><div><hr></div><h2><strong>Get in touch</strong></h2><p>You can find me on <a href="https://www.linkedin.com/in/patilamrut/">LinkedIn</a> or <a href="https://twitter.com/realamrutpatil">X</a>.</p><p>If you would like to request a topic to read, please feel free to contact me directly via LinkedIn or X.</p>]]></content:encoded></item><item><title><![CDATA[TCP #129: The scoring framework that picks ECS, EKS, or Lambda.]]></title><description><![CDATA[A weighted decision matrix across team maturity, operations, compliance, cost, and ownership with worked examples.]]></description><link>https://www.thecloudplaybook.com/p/ecs-eks-lambda-decision-framework</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/ecs-eks-lambda-decision-framework</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Thu, 02 Jul 2026 15:01:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Zdys!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a874e54-0a7f-4b07-b9bb-da4f9fe1e68e_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most AWS compute decisions are made in a conference room. Two engineers argue for their preferred option. A staff engineer summarizes. The team picks. The decision is filed.</p><p>The conversation produces an outcome. It does not produce a defensible record. Six months later, when a similar workload arrives, the next team has the same conversation. There is no organizational memory. There is no consistency across teams. The compute estate fragments.</p><p>A scoring framework solves this. The framework does not eliminate judgment. It structures judgment so that the same workload produces the same recommendation regardless of which engineer scores it. It also produces a written record that future teams can reference and that leadership can audit when the cost trajectory deviates from the forecast.</p><p>The framework below is the one I would use with a real engineering team. It is calibrated for SaaS workloads, weighted for the dimensions that matter at 18-month horizons, and written at a level of detail that can be applied without further interpretation.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Zdys!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a874e54-0a7f-4b07-b9bb-da4f9fe1e68e_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Zdys!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a874e54-0a7f-4b07-b9bb-da4f9fe1e68e_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!Zdys!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a874e54-0a7f-4b07-b9bb-da4f9fe1e68e_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!Zdys!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a874e54-0a7f-4b07-b9bb-da4f9fe1e68e_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!Zdys!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a874e54-0a7f-4b07-b9bb-da4f9fe1e68e_1448x1086.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Zdys!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a874e54-0a7f-4b07-b9bb-da4f9fe1e68e_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7a874e54-0a7f-4b07-b9bb-da4f9fe1e68e_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1559000,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/203907995?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a874e54-0a7f-4b07-b9bb-da4f9fe1e68e_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Zdys!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a874e54-0a7f-4b07-b9bb-da4f9fe1e68e_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!Zdys!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a874e54-0a7f-4b07-b9bb-da4f9fe1e68e_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!Zdys!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a874e54-0a7f-4b07-b9bb-da4f9fe1e68e_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!Zdys!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a874e54-0a7f-4b07-b9bb-da4f9fe1e68e_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Why Most Frameworks Fail</h2><p>Frameworks fail in three predictable ways.</p><ol><li><p><strong>Too many dimensions.</strong> The team scores 12 attributes. The scoring takes three hours. The team abandons the framework after the second use because it feels like overhead.</p></li><li><p><strong>Wrong dimensions.</strong> The framework scores attributes that do not predict success. Language support, AWS region availability, and console UX. These attributes feel important, but rarely change the recommendation.</p></li><li><p><strong>No threshold rules.</strong> The framework produces a score for each option, but no rule for what the score means. Three options score within 10 percent of each other. The team argues again. The framework produced no decision support.</p></li></ol><p>The framework below is calibrated against these failure modes. Five dimensions, weighted by 18-month impact, with explicit threshold rules that produce a recommendation rather than a debate.</p>
      <p>
          <a href="https://www.thecloudplaybook.com/p/ecs-eks-lambda-decision-framework">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[TCP #128: ECS, EKS, and Lambda are not the same decision]]></title><description><![CDATA[How team maturity, operational burden, and tenancy reshape the AWS compute choice, and what teams get wrong by treating them as alternatives]]></description><link>https://www.thecloudplaybook.com/p/ecs-eks-lambda-not-interchangeable</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/ecs-eks-lambda-not-interchangeable</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Sun, 28 Jun 2026 14:21:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!CV9T!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9744ee-78a5-41c3-ac12-5d3b1da2e3da_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most engineering teams treat the choice of AWS compute as a matter of preference. </p><p>Lambda for the team that wants serverless. </p><p>ECS for the team that prefers containers. </p><p>EKS for the team that wants Kubernetes. </p><p>The conversation runs for an hour, the team picks the option that fits how it already works, and the service ships.</p><p>Eighteen months later, the choice is the most consequential architectural commitment the team has made. It governs cost, on-call load, hiring, audit posture, deployment cadence, and the team&#8217;s ability to take on the next workload type. It is harder to change than the database choice and almost as hard as changing the account topology.</p><p>The teams that get this right do not pick based on preference. </p><p>They pick based on five dimensions that determine whether the choice will compound value or compound debt: <strong>team maturity, operational burden, latency profile, tenancy model, and compliance exposure.</strong></p><p>The teams that get this wrong pick based on what they already know.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CV9T!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9744ee-78a5-41c3-ac12-5d3b1da2e3da_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CV9T!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9744ee-78a5-41c3-ac12-5d3b1da2e3da_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!CV9T!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9744ee-78a5-41c3-ac12-5d3b1da2e3da_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!CV9T!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9744ee-78a5-41c3-ac12-5d3b1da2e3da_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!CV9T!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9744ee-78a5-41c3-ac12-5d3b1da2e3da_1448x1086.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CV9T!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9744ee-78a5-41c3-ac12-5d3b1da2e3da_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5f9744ee-78a5-41c3-ac12-5d3b1da2e3da_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2027535,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/202943599?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9744ee-78a5-41c3-ac12-5d3b1da2e3da_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CV9T!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9744ee-78a5-41c3-ac12-5d3b1da2e3da_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!CV9T!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9744ee-78a5-41c3-ac12-5d3b1da2e3da_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!CV9T!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9744ee-78a5-41c3-ac12-5d3b1da2e3da_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!CV9T!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9744ee-78a5-41c3-ac12-5d3b1da2e3da_1448x1086.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Why Familiarity Is the Wrong Decision Criterion</h2><p>Familiarity is the default tiebreaker, and it is the wrong one.</p><p>The argument for familiarity is operational. The team that picks the technology it knows ships faster, debugs faster, and produces fewer incidents in the first six months. This is correct. It is also short-sighted.</p><p>The argument against familiarity is structural. Most engineering teams will make compute choices for workloads that grow in volume, expand in tenancy, and acquire new compliance scope. Familiarity at year zero produces operational comfort. By year two, the workload no longer matches the compute model the team chose, and the cost of the mismatch is paid by every team that runs on the same platform.</p><p>The teams that pick on familiarity end up with three compute platforms five years later, each one chosen by a different team for a different workload, each one operationally expensive. The platform team inherits the burden of supporting all three.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Cloud Playbook is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>The Five Dimensions That Should Drive the Choice</h2><p>Compute selection is a multi-dimensional decision. The five dimensions below yield recommendations that differ from those of the familiarity heuristic.</p><ol><li><p><strong>Team maturity</strong></p><p>Lambda is the lowest operational burden of the three. ECS is moderate. EKS is high. The team&#8217;s operational maturity does not need to match the platform&#8217;s complexity at year zero, but it needs to match within 18 months. A team with no Kubernetes experience that picks EKS will be operating an architecture they cannot debug for two quarters. A team with deep container experience that picks Lambda will hit cold-start, package size, and orchestration limits within a year.</p></li><li><p><strong>Operational burden</strong></p><p>Lambda offloads scaling, patching, and runtime management to AWS. ECS Fargate offloads node management, but the team still owns container lifecycle. EKS carries almost all of the operational burden with the team. The right question is not &#8220;which is easier.&#8221; The right question is &#8220;what fraction of the team&#8217;s capacity can be dedicated to running this platform six months from now?&#8221; If the answer is 0 percent, then the choice is between Lambda and Fargate. If the answer is 30 percent, EKS is on the table.</p></li><li><p><strong>Latency profile</strong></p><p>Lambda has cold-start variance, which can be addressed but never eliminated. The variance is incompatible with sub-100ms P99 SLOs for low-volume services. ECS and EKS have warm capacity by default. For latency-critical workloads, the cold-start question is not theoretical. It is the deciding factor.</p></li><li><p><strong>Tenancy model</strong></p><p>Lambda is shared-tenant by default, and tenant isolation is operationally awkward. ECS and EKS support tenant isolation through cluster topology, namespace boundaries, and IAM scope. For SaaS workloads that will need per-tenant isolation in the future, the tenancy commitment made today shapes whether that future is achievable.</p></li><li><p><strong>Compliance exposure</strong></p><p>All three services support common compliance frameworks. The differences are in the operational effort required to produce evidence. Lambda&#8217;s logging and configuration boundary is the simplest. EKS produces the most evidence but also requires the most work to properly configure it. The compliance frameworks the team will face in the next 24 months should drive this dimension.</p></li></ol><div><hr></div><h2>Why Lambda Is Misused</h2><p>Lambda is the most misused compute service in AWS. The misuse is not technical. It is economic.</p><p>Lambda is priced per invocation and per millisecond of execution. At low volume, this is dramatically cheaper than running a container fleet. At high volume, it is dramatically more expensive. The cost crossover varies by workload, but for typical request-response services, it sits somewhere between 100 and 500 RPS sustained. Below the crossover, Lambda saves money. Above the crossover, Lambda costs 2 to 5 times as much as an equivalent ECS or EKS deployment.</p><p>Most teams choose Lambda based on their current volume. They forecast linearly. They do not model the workload at the volume it will reach in 18 months. By the time the cost trajectory becomes visible in the AWS bill, the architecture is load-bearing, and the migration cost is high.</p><p>The Lambda choice is not wrong. The Lambda choice without a written cost trajectory at 10x current volume is wrong.</p><div><hr></div><h2>Why EKS Is Often the Wrong Default</h2><p>EKS has become the default container choice for engineering teams that want maximum flexibility. The argument is reasonable: Kubernetes is the broadest ecosystem, the most portable, the most powerful.</p><p>The argument also misses two structural realities.</p><p>First, the operational burden is not optional. EKS requires the team to operate cluster networking, ingress, pod scheduling, node lifecycle, secrets management, observability instrumentation, and security boundaries. AWS manages the control plane. The team manages everything else. The team that picks EKS is committing to roughly 0.5 to 1.0 FTE of platform work that does not directly produce product value.</p><p>Second, the flexibility is rarely used. Most workloads that pick EKS do not use the features that justify the operational burden. They run a deployment, a service, and an ingress. ECS Fargate would run the same workload with a fraction of the operational overhead and no portability cost; the team would actually exercise.</p><p>EKS is the right choice when the team has a portfolio of workloads that justify the platform investment, a strong commitment to Kubernetes patterns across services, or a specific feature, such as custom schedulers, that ECS does not support. It is the wrong choice for the team to pick it because the rest of the industry is using it.</p><div><hr></div><h2>Why ECS Is Underestimated</h2><p>ECS does not have the brand prestige of Kubernetes. It lacks the elegance of Lambda. It has the property that matters more than either: it requires the least operational burden for the broadest range of workloads.</p><p>ECS Fargate runs containers without node management. It integrates with the rest of AWS through native primitives rather than abstractions. It produces logs, metrics, and traces through standard AWS instrumentation. It has been generally available longer than EKS and is operationally mature.</p><p>For most SaaS workloads, ECS Fargate is the correct default. It supports latency profiles, tenancy models, and compliance frameworks. It does not require the team to operate a Kubernetes cluster. It does not lock the team into Lambda&#8217;s pricing trajectory.</p><p>The teams that build on ECS as the default and reach for Lambda or EKS only when the workload demands it produce platforms with lower operational burden and better cost trajectories than the teams that pick on preference.</p><div><hr></div><h2>Four Questions That Drive the Right Choice</h2><p>Before the team picks a compute service, four questions should be answered in writing. The questions take an hour. They prevent the wrong choice from compounding for years.</p><ol><li><p><strong>What is the team&#8217;s operational capacity for this workload at month 12</strong></p><p>Not month zero. The team has Slack at month zero because the workload is new. By month 12, the workload is one of several. If the team will not have 30 percent of an FTE available to run a Kubernetes cluster, EKS is not on the table.</p></li><li><p><strong>At 10x current volume, what does this cost?</strong></p><p>Lambda&#8217;s cost trajectory hides at low volume and surfaces at high volume. Modeling 10x growth at decision time exposes the cost cliff. If the model shows Lambda becoming uneconomic at the forecast volume, the choice should be ECS or EKS.</p></li><li><p><strong>What latency SLO does the workload need?</strong></p><p>Sub-100ms P99 with low volume is incompatible with Lambda cold-start variance unless provisioned concurrency is paid for. For workloads with strict latency requirements and low base load, ECS or EKS are the only viable choices.</p></li><li><p><strong>What tenancy and compliance requirements will the workload need in 24 months?</strong> </p><p>If the future includes per-tenant isolation for regulated customers, the compute choice today should not foreclose that future. Lambda&#8217;s tenancy model is the most awkward to evolve. ECS and EKS both support evolution but require different paths.</p></li></ol><p>The team that answers these four questions makes a different choice than the team that picks based on preference. The choice is not always the same. It is always more defensible.</p><div><hr></div><h2>What Changes When Computing Is a Risk Decision</h2><p>Engineering organizations that treat compute selection as a risk decision rather than a preference produce different outcomes.</p><p>The platform team consolidates on fewer compute models, reducing operational burden and improving the team&#8217;s ability to ship golden paths.</p><p>The cost forecast is accurate because each workload was modeled during growth, before the architecture shipped.</p><p>The compliance posture is consistent because the compute choice was filtered through the compliance scope at decision time.</p><p>The team&#8217;s hiring profile becomes coherent because computed commitments determine which kinds of engineers the team needs to recruit, develop, and retain.</p><p>The choice of a computer does not become invisible. It becomes the single most important architectural commitment the team makes. The teams that treat it that way ship platforms that compound. The teams that treat it as a preference ship platforms that fragment.</p><div><hr></div><p>On Wednesday, paid subscribers get the complete <strong>ECS vs EKS vs Lambda decision framework</strong>: a scoring matrix that evaluates a candidate workload across team maturity, operational complexity, compliance impact, cost trajectory, and ownership. </p><p>The framework includes weighted scoring, the threshold rules that drive the recommendation, and three worked examples covering an internal service, a customer-facing API, and a regulated multi-tenant workload.</p><div><hr></div><p><strong>Upgrade If You Need Implementation, Not Just Ideas</strong></p><p>If you&#8217;re using these emails to guide real decisions on your platform, you&#8217;ll get more leverage from the paid version of The Cloud Playbook.</p><p>The free newsletter gives you patterns and language.</p><p>The paid newsletter turns those patterns into implementation kits you can ship inside a quarter:</p><ul><li><p>Concrete rollout plans (90&#8209;day roadmaps for each pattern)</p></li><li><p>Templates and checklists (policies, runbooks, tagging schemes, review checklists)</p></li><li><p>Real examples from high&#8209;stakes AWS environments (what we actually shipped and why)</p></li></ul><p>If the paid side doesn&#8217;t save you more than the subscription in <strong>one</strong> incident, audit cycle, or bad migration you avoid, you should cancel and keep the playbooks.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe&quot;,&quot;text&quot;:&quot;Upgrade to the Paid Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thecloudplaybook.com/subscribe"><span>Upgrade to the Paid Cloud Playbook</span></a></p><div><hr></div><h2><strong>That&#8217;s it for today!</strong></h2><p>Did you enjoy this newsletter issue?</p><p>Share with your friends, colleagues, and your favorite social media platform.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Cloud Playbook</span></a></p><p><strong>Until next week &#8212; Amrut</strong></p><div><hr></div><h2><strong>Get in touch</strong></h2><p>You can find me on <a href="https://www.linkedin.com/in/patilamrut/">LinkedIn</a> or <a href="https://twitter.com/realamrutpatil">X</a>.</p><p>If you would like to request a reading topic, please feel free to contact me directly via LinkedIn or X.</p>]]></content:encoded></item><item><title><![CDATA[TCP #127: The ADR template that turns architecture into a written contract]]></title><description><![CDATA[Decision context, tradeoffs, risk, cost, compliance, and rollback captured in a format leadership can read.]]></description><link>https://www.thecloudplaybook.com/p/aws-architecture-decision-record-template</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/aws-architecture-decision-record-template</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Sun, 21 Jun 2026 14:25:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!PrZx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4236d76b-51e0-4a38-954e-716a7273586f_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Architecture Decision Records are not new. </p><p>The format has been documented for over a decade. Most engineering organizations have adopted it. Most engineering organizations also continue to ship architectures that look fine on day one and become liabilities by month 18.</p><p>The ADR is not the problem. The way most ADRs are written is.</p><p>Most ADR templates capture three things: <strong>context, decision, consequences.</strong> </p><p>The format is clean, the discipline is real, but the output is rarely actionable. </p><p>The &#8220;consequences&#8221; section is a paragraph of speculation. The decision is recorded. The risks are not modeled. The architecture ships, the document is filed, and the next team to touch the system has no instrument for evaluating whether the original decision still holds.</p><p>A well-designed ADR template fixes this. </p><p>It forces the team to write down what would change the decision, what the cost trajectory looks like, what compliance scope is closed off, and what the rollback path is. </p><p>The template does not produce better decisions because the team is smarter. It produces better decisions because the format will not accept the easy answers.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PrZx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4236d76b-51e0-4a38-954e-716a7273586f_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PrZx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4236d76b-51e0-4a38-954e-716a7273586f_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!PrZx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4236d76b-51e0-4a38-954e-716a7273586f_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!PrZx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4236d76b-51e0-4a38-954e-716a7273586f_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!PrZx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4236d76b-51e0-4a38-954e-716a7273586f_1448x1086.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PrZx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4236d76b-51e0-4a38-954e-716a7273586f_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4236d76b-51e0-4a38-954e-716a7273586f_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2191665,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/202940946?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4236d76b-51e0-4a38-954e-716a7273586f_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PrZx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4236d76b-51e0-4a38-954e-716a7273586f_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!PrZx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4236d76b-51e0-4a38-954e-716a7273586f_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!PrZx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4236d76b-51e0-4a38-954e-716a7273586f_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!PrZx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4236d76b-51e0-4a38-954e-716a7273586f_1448x1086.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>What an ADR Is For</h2><p>An ADR is a written contract between the team that makes a decision and the team that will inherit it.</p><p>The contract states: this is what we knew, this is what we evaluated, this is what we chose, this is what we accepted. The future team reads it and knows whether the conditions that justified the decision still hold. If they do, the architecture continues. If they do not, the team has the documentation to argue for change.</p><p>An ADR is not a justification document. It is not a sales pitch for the decision the team has already made. The most useful ADRs document the alternatives the team rejected with the same rigor as the alternative the team accepted. A reader who cannot tell from the ADR why the second-place option lost is reading a document that failed to do its job.</p><p>An ADR is also not a one-time artifact. It has a lifecycle: proposed, accepted, deprecated, superseded. A team that accepts an ADR commits to revisiting it on a defined cadence and updating its status when the conditions change.</p><p>The discipline of writing an ADR slowly compounds organizational memory. Decisions that would have been forgotten and re-debated three years later instead become the foundation for the next decision.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Cloud Playbook is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>The Eleven-Section ADR Template</h2><p>The template below is ten percent more structured than the canonical Michael Nygard ADR format. The additional structure is what produces the behavioral change.</p><p><strong>1. Title</strong> </p><p><strong>2. Status</strong> </p><p><strong>3. Context</strong> </p><p><strong>4. Decision</strong> </p><p><strong>5. Alternatives Considered</strong> </p><p><strong>6. Operational Risk</strong> </p><p><strong>7. Cost Trajectory</strong> </p><p><strong>8. Compliance and Tenancy Impact</strong> </p><p><strong>9. Rollback Path</strong> </p><p><strong>10. Review Cadence</strong> </p><p><strong>11. Related ADRs</strong></p><p>Each section has a defined purpose, a length target, and a quality bar. ADRs that fail to meet the quality bar are returned for revision before acceptance.</p><div><hr></div><h2>Section-by-Section Specifications</h2><p><strong>1. Title.</strong> A complete sentence stating the decision in the active voice. Not &#8220;Compute Choice for Service X.&#8221; Instead: &#8220;Service X will run on ECS Fargate rather than Lambda or EKS.&#8221; The title alone should communicate what was decided.</p><p><strong>2. Status.</strong> One of: Proposed, Accepted, Deprecated, Superseded by ADR-NNN. Status is dated. A Proposed ADR with no status change in 14 days is escalated to the architecture review forum.</p><p><strong>3. Context.</strong> Two to four paragraphs. Why is the decision being made now? What problem is being solved? What constraints (deadline, budget, customer commitment, compliance scope) are operating? The reader must be able to determine from the context section whether the same decision would be made today, regardless of what the decision was.</p><p><strong>4. Decision.</strong> Two paragraphs maximum. What is being decided? Specific service names, specific configurations, specific topology. Avoid abstraction. A decision that reads &#8220;we will use a managed service for queueing&#8221; is not a decision. A decision that reads &#8220;we will use SQS Standard with a 14-day retention and a dead-letter queue redrive policy of 5 attempts&#8221; is.</p><p><strong>5. Alternatives Considered.</strong> Three or more alternatives. Each alternative gets a paragraph. The paragraph names the alternative, names two reasons it could have been the right choice, and names the specific reason it was not. Do not write &#8220;we considered X but it did not fit.&#8221; Write &#8220;we considered EKS because three teams already operate it and the operational tooling exists, but rejected it because the workload is single-tenant and bursty and would require a dedicated node group with low utilization.&#8221;</p><p><strong>6. Operational Risk.</strong> Three paragraphs. The first names the most likely failure mode the architecture introduces and the recovery path. The second states the recovery-time objective the architecture supports. The third names the operational burden the team accepts: on-call obligations, required runbooks, and required monitoring instrumentation.</p><p><strong>7. Cost Trajectory.</strong> A table with three rows: cost at current workload, cost at 10x current workload, cost at the volume the team forecasts in 24 months. Each row shows the monthly figure and the assumption set. Below the table, one paragraph naming the cost cliff (the point at which a different architecture would be cheaper) and the conditions that would trigger reconsideration.</p><p><strong>8. Compliance and Tenancy Impact.</strong> Three paragraphs. The first names the compliance frameworks the architecture supports today (SOC 2, FedRAMP, HIPAA, ISO 27001, PCI). The second names the compliance frameworks or customer types the architecture excludes or makes harder. The third names the tenancy model the architecture commits to (shared, isolated, hybrid) and what would change if that commitment is revisited.</p><p><strong>9. Rollback Path.</strong> Two paragraphs. The first names the conditions under which this decision would be reversed. The second names the cost, duration, and team capacity required to reverse it. The rollback path is the most important section to write honestly. A decision with no rollback path is not a decision. It is a one-way door.</p><p><strong>10. Review Cadence.</strong> A single sentence stating the date the ADR will be reviewed and the trigger for an earlier review. Default cadence: 12 months from acceptance. Earlier triggers: cost trajectory variance above 20 percent, an incident attributable to the decision, a customer or compliance scope change.</p><p><strong>11. Related ADRs.</strong> A list of ADR numbers with one-line annotations describing the relationship: depends on, supersedes, related to, conflicts with. This section is what produces organizational memory. Future teams use the related ADRs to navigate the decision history.</p><div><hr></div><h2>Worked Example: Compute Selection for a New Service</h2><p><strong>Title:</strong> The customer-events ingestion service will run on ECS Fargate rather than Lambda or EKS.</p><p><strong>Status:</strong> Accepted, 2026-06-10.</p><p><strong>Context:</strong> The customer-events service ingests webhook events from third-party integrations and writes them to an internal event bus. Forecast volume is 200 RPS sustained, 2,000 RPS burst, with an average payload size of 4KB. The service is expected to be a foundation for customer-facing event replay capabilities, which will increase volume by an estimated 5x within 12 months. The team operating this service has two engineers familiar with ECS and one engineer familiar with Lambda. EKS is operated by the platform team, but no application teams currently deploy to it.</p><p><strong>Decision:</strong> The service runs on ECS Fargate with a minimum of 2 tasks and an autoscaling target of 60% CPU utilization, scaling up to a maximum of 20 tasks. The service is deployed to the platform team&#8217;s standard ECS cluster in production with the existing observability stack. The container image is built on the standard Python 3.12 base image and includes the platform team&#8217;s standard observability sidecar.</p><p><strong>Alternatives Considered:</strong></p><p>Lambda was considered for the speed-of-shipping advantage and the lack of cluster management overhead. Two engineers had already prototyped the workload on Lambda. Rejected because at the 24-month forecast volume of 1,000 RPS sustained, Lambda costs are 3.4x ECS costs based on internal calculator output, and the cold-start variance was incompatible with the latency SLO of P99 under 200ms.</p><p>EKS was considered for consistency with the platform team&#8217;s longer-term direction and access to the broader Kubernetes ecosystem. Rejected because the team operating this service has no Kubernetes experience and the workload does not justify the operational learning curve. The platform team confirmed that ECS is supported indefinitely as a first-class option.</p><p>A self-managed EC2 fleet was considered briefly for cost predictability at high volume. Rejected because the operational burden of patching, scaling, and logging would consume engineering time disproportionate to the cost savings.</p><p><strong>Operational Risk:</strong> The most likely failure mode is autoscaling lag during sudden traffic spikes from third-party integration partners. The recovery path is to increase the autoscaling minimum capacity manually via the ECS console; runbook RB-042 documents the procedure. The recovery time objective is 5 minutes for autoscaling-related capacity issues. The team accepts on-call rotation for this service, integration into the existing PagerDuty schedule, and the maintenance of two runbooks.</p><p><strong>Cost Trajectory:</strong></p><p><strong>Current Load (200 RPS)</strong></p><p><strong>Estimated Monthly Cost:</strong> $480</p><p><strong>Assumption:</strong> Average of 4 ECS Fargate tasks running in us-east-1 using the current task configuration.</p><div><hr></div><p><strong>24-Month Forecast (1,000 RPS)</strong></p><p><strong>Estimated Monthly Cost:</strong> $2,200</p><p><strong>Assumption:</strong> Average of 9 ECS Fargate tasks running with the same task sizing and configuration.</p><div><hr></div><p><strong>10&#215; Growth Scenario (2,000 RPS)</strong></p><p><strong>Estimated Monthly Cost:</strong> $4,200</p><p><strong>Assumption:</strong> Average of 18 ECS Fargate tasks running with the same task sizing and configuration.</p><p>The cost cliff is at approximately 4,000 RPS sustained, where moving to ECS on EC2 with reserved instances would produce a 35 percent saving. The trigger for reconsideration is sustained volume above 3,500 RPS for two consecutive months.</p><p><strong>Compliance and Tenancy Impact:</strong> This architecture supports SOC 2 Type II without modification using the platform team&#8217;s standard logging and access controls. FedRAMP Moderate is supported by the addition of the GovCloud deployment pattern, as documented in ADR-039. The architecture is shared-tenant by default; isolating a specific customer to a dedicated task fleet is a 1-week effort if required.</p><p><strong>Rollback Path:</strong> Reversing this decision to Lambda would require approximately 3 weeks of engineering effort (rewriting the deployment pipeline, refactoring the service initialization, and recreating the observability instrumentation). Reversing to EKS would require approximately 6 weeks, including team training. Either rollback would require a coordinated cutover with the third-party integration partners. The rollback decision would be triggered by a sustained variance in cost trajectory of more than 30 percent or by a platform-level deprecation of ECS Fargate.</p><p><strong>Review Cadence:</strong> This ADR will be reviewed on 2027-06-10. Earlier review triggers: sustained volume above 3,500 RPS, cost variance above 20 percent, FedRAMP customer commitment.</p><p><strong>Related ADRs:</strong> ADR-014 (platform observability standard, depends on), ADR-027 (ECS cluster topology, depends on), ADR-039 (GovCloud deployment pattern, related to).</p><div><hr></div><h2>How to Roll Out the ADR Template</h2><p>Adopting this template across a team is a behavioral change. Three rollout steps reduce friction.</p><ul><li><p><strong>Start with one new decision.</strong> Do not retroactively author ADRs for existing systems. Pick the next non-trivial architecture decision the team faces and write it in the new format. Have the team review the ADR before approving the decision. The team learns the format by using it on a real decision under real pressure.</p></li><li><p><strong>Treat the first three ADRs as calibration.</strong> The first ADR will be too long. The second will be too short. The third will start to feel right. Resist the temptation to fix the first two. The format is the point; the calibration is automatic.</p></li><li><p><strong>Make ADRs visible to leadership.</strong> ADRs are not an internal team artifact. They are how the platform team communicates architectural commitment to the rest of the engineering organization. Publish accepted ADRs to a known location and reference them in MBR documentation. Leadership that can see the ADRs makes better resource decisions about the platform team.</p></li></ul><div><hr></div><h2>What to Measure to Evaluate the ADR Practice</h2><p>The ADR practice is itself an investment. It should be measured.</p><ul><li><p><strong>Time-to-acceptance for new ADRs.</strong> From Proposed to Accepted, the target is 7 business days. Longer than that indicates either the decision is not ready to be made or the review forum is not meeting often enough.</p></li><li><p><strong>Percentage of architecture decisions documented as ADRs.</strong> Track this manually for the first quarter. Target: above 80 percent of decisions touching shared infrastructure, compliance scope, or production tenancy. Below 60 percent indicates the format is being seen as overhead and rolled around.</p></li><li><p><strong>Frequency of ADR reference in subsequent decisions.</strong> Count how often new ADRs cite prior ADRs in the Related ADRs section. This is the leading indicator that organizational memory is compounding. Target: at least 60 percent of new ADRs reference at least one prior ADR by month 6.</p></li><li><p><strong>Outcome accuracy.</strong> When an ADR comes up for review, evaluate how well the original cost trajectory, operational risk, and compliance impact predictions matched reality. Predictions that miss by more than 30 percent indicate the team is not modeling rigorously enough at decision time.</p></li></ul><p>Review these metrics in the platform's monthly operating review. The ADR practice that is not measured is the practice that will be abandoned in the next quarter under deadline pressure.</p><div><hr></div><h2>When ADRs Become the Architecture Conversation</h2><p>A mature ADR practice produces a specific change in how the engineering organization talks about architecture.</p><p>Discussions move from speculation to documents. &#8220;What did we decide about X?&#8221; becomes &#8220;let me pull up ADR-073.&#8221; </p><p>Disagreements become more productive: the team is arguing about the document, not about memories. New engineers ramp faster: the architecture has a written history, and the history reads as a sequence of decisions rather than a state of the world.</p><p>Most importantly, the architecture stops accumulating decisions that nobody can defend. The ADR forces every load-bearing choice through a format that requires defense at the moment of the choice. Choices that cannot be defended in the format are reconsidered before they ship.</p><div><hr></div><h2>Upgrade If You Need Implementation, Not Just Ideas</h2><p>If you&#8217;re using these emails to guide real decisions on your platform, you&#8217;ll get more leverage from the paid version of The Cloud Playbook.</p><p>The free newsletter gives you patterns and language.</p><p>The paid newsletter turns those patterns into implementation kits you can ship inside a quarter:</p><ul><li><p>Concrete rollout plans (90&#8209;day roadmaps for each pattern)</p></li><li><p>Templates and checklists (policies, runbooks, tagging schemes, review checklists)</p></li><li><p>Real examples from high&#8209;stakes AWS environments (what we actually shipped and why)</p></li></ul><p>If the paid side doesn&#8217;t save you more than the subscription in <strong>one</strong> incident, audit cycle, or bad migration you avoid, you should cancel and keep the playbooks.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe&quot;,&quot;text&quot;:&quot;Upgrade to the Paid Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thecloudplaybook.com/subscribe"><span>Upgrade to the Paid Cloud Playbook</span></a></p><div><hr></div><h2><strong>That&#8217;s it for today!</strong></h2><p>Did you enjoy this newsletter issue?</p><p>Share with your friends, colleagues, and your favorite social media platform.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Cloud Playbook</span></a></p><p><strong>Until next week &#8212; Amrut</strong></p><div><hr></div><h2><strong>Get in touch</strong></h2><p>You can find me on <a href="https://www.linkedin.com/in/patilamrut/">LinkedIn</a> or <a href="https://twitter.com/realamrutpatil">X</a>.</p><p>If you would like to request a topic to read, please feel free to contact me directly via LinkedIn or X.</p><p></p>]]></content:encoded></item><item><title><![CDATA[TCP #126: Architecture decisions are risk decisions. Most teams forget that.]]></title><description><![CDATA[How to evaluate AWS architecture choices on what they cost you in 18 months, not what they save you this sprint.]]></description><link>https://www.thecloudplaybook.com/p/aws-architecture-decisions-reduce-risk</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/aws-architecture-decisions-reduce-risk</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Thu, 18 Jun 2026 15:05:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!1Wn7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f94f93a-b4c5-4aa0-bb51-a4a292829afb_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Every AWS architecture decision is made under deadline pressure. The product team needs the feature. The team picks the design that ships fastest. The work goes out, the feature performs, and the decision is logged as a success.</p><p>Eighteen months later, the same architecture is the reason a migration is delayed, an audit stalls, an incident takes four hours to recover from, or a tenant cannot be onboarded without manual intervention.</p><p>The decision was not wrong on the day it was made. It was made against the wrong frame.</p><p>Architecture choices are not feature decisions. They are risk decisions. The team that treats them as feature decisions ships faster in the short term and pays interest forever. The team that treats them as risk decisions ships almost as fast and stops paying interest after the first iteration.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1Wn7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f94f93a-b4c5-4aa0-bb51-a4a292829afb_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1Wn7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f94f93a-b4c5-4aa0-bb51-a4a292829afb_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!1Wn7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f94f93a-b4c5-4aa0-bb51-a4a292829afb_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!1Wn7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f94f93a-b4c5-4aa0-bb51-a4a292829afb_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!1Wn7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f94f93a-b4c5-4aa0-bb51-a4a292829afb_1448x1086.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1Wn7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f94f93a-b4c5-4aa0-bb51-a4a292829afb_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6f94f93a-b4c5-4aa0-bb51-a4a292829afb_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1755484,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/202048999?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f94f93a-b4c5-4aa0-bb51-a4a292829afb_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1Wn7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f94f93a-b4c5-4aa0-bb51-a4a292829afb_1448x1086.png 424w, https://substackcdn.com/image/fetch/$s_!1Wn7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f94f93a-b4c5-4aa0-bb51-a4a292829afb_1448x1086.png 848w, https://substackcdn.com/image/fetch/$s_!1Wn7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f94f93a-b4c5-4aa0-bb51-a4a292829afb_1448x1086.png 1272w, https://substackcdn.com/image/fetch/$s_!1Wn7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f94f93a-b4c5-4aa0-bb51-a4a292829afb_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Why the Risk Frame Gets Lost</h2><p>Engineering organizations frame architecture as a delivery problem because delivery is the visible bottleneck.</p><p>The roadmap is fixed. The deadline is fixed. </p><p>The team is asked: Which AWS service can we use to ship this on time? The conversation is dominated by what the design enables today, not what it constrains tomorrow.</p><p>This produces predictable failure patterns. The team picks Lambda for a workload that grows into a six-figure monthly bill. The team picks DynamoDB for a workload that turns out to need transactional integrity across tables. The team selects a single-account deployment, which prevents a future move to a multi-tenant platform. The team picks a managed service in a region that does not meet the next compliance scope.</p><p>None of these decisions is bad in isolation. Each one solves the immediate problem. The cost is not visible at the time of decision. It surfaces during an event the team did not plan for: a migration, an incident, an audit, a customer in a regulated industry, or a competitor that built the same product more cheaply.</p><p>By that point, the architecture is load-bearing. Changing it is expensive enough that the team continues to pay the interest rather than refinance.</p><div><hr></div><h2>The Three Categories of Future Risk</h2><p>Every AWS architecture decision implicates three categories of future risk. The team that names them at decision time makes different choices than the team that does not.</p>
      <p>
          <a href="https://www.thecloudplaybook.com/p/aws-architecture-decisions-reduce-risk">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[TCP #125: The MBR template that changes how leadership funds platform.]]></title><description><![CDATA[Delivery, reliability, cost, security, and risk, formatted for the engineering leadership audience.]]></description><link>https://www.thecloudplaybook.com/p/platform-engineering-mbr-template</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/platform-engineering-mbr-template</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Sun, 14 Jun 2026 14:30:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3sGw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7687625f-2bdd-4f72-9c3e-2ef7c4060c80_1086x1448.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Platform engineering leaders send monthly updates. Leadership reads them. Nothing changes.</p><p>The headcount conversation in the next planning cycle plays out the same way it did the previous quarter. The roadmap discussion still treats the platform as overhead. The budget conversation defaults to &#8220;do more with what you have.&#8221;</p><p>The problem is rarely effort. Most platform leaders write thorough updates. </p><p>The problem is the format. The update is structured around what the team did, not around what the business should care about. Leadership reads it, finds nothing they can act on, and moves on.</p><p>A Monthly Business Review designed for executive consumption fixes this. Not by inflating activity. By restructuring the same information so leadership can see the value, the risk, and the investment case in a format they already know how to read.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3sGw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7687625f-2bdd-4f72-9c3e-2ef7c4060c80_1086x1448.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3sGw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7687625f-2bdd-4f72-9c3e-2ef7c4060c80_1086x1448.png 424w, https://substackcdn.com/image/fetch/$s_!3sGw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7687625f-2bdd-4f72-9c3e-2ef7c4060c80_1086x1448.png 848w, https://substackcdn.com/image/fetch/$s_!3sGw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7687625f-2bdd-4f72-9c3e-2ef7c4060c80_1086x1448.png 1272w, https://substackcdn.com/image/fetch/$s_!3sGw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7687625f-2bdd-4f72-9c3e-2ef7c4060c80_1086x1448.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3sGw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7687625f-2bdd-4f72-9c3e-2ef7c4060c80_1086x1448.png" width="727.9951171875" height="970.66015625" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7687625f-2bdd-4f72-9c3e-2ef7c4060c80_1086x1448.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1448,&quot;width&quot;:1086,&quot;resizeWidth&quot;:727.9951171875,&quot;bytes&quot;:1762489,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/201888041?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7687625f-2bdd-4f72-9c3e-2ef7c4060c80_1086x1448.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3sGw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7687625f-2bdd-4f72-9c3e-2ef7c4060c80_1086x1448.png 424w, https://substackcdn.com/image/fetch/$s_!3sGw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7687625f-2bdd-4f72-9c3e-2ef7c4060c80_1086x1448.png 848w, https://substackcdn.com/image/fetch/$s_!3sGw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7687625f-2bdd-4f72-9c3e-2ef7c4060c80_1086x1448.png 1272w, https://substackcdn.com/image/fetch/$s_!3sGw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7687625f-2bdd-4f72-9c3e-2ef7c4060c80_1086x1448.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>What an MBR Is and Is Not</h2><p>A Monthly Business Review is a 4 to 6 page document delivered on a fixed cadence to a defined audience: VP of Engineering, CTO, and any executive whose budget the platform team draws from.</p><p>It is not a status update. Status updates report activity. MBRs report outcomes, decisions, and asks.</p><p>It is not a retrospective. Retrospectives are internal. MBRs are external-facing artifacts designed to inform investment decisions and surface risks that need leadership attention.</p><p>It is not a slide deck. The MBR is a written document that a CTO can read in 15 minutes without you in the room. If it requires you to walk leadership through it, the format is wrong.</p><p>The defining property of an MBR is asymmetry of effort. The platform leader spends three hours writing it. Leadership spends 15 minutes reading it. That trade is the format&#8217;s entire value.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Cloud Playbook is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>The Six-Section Structure</h2><p>Every MBR follows the same six sections in the same order. Consistency is the point. </p><p>Leadership learns the format once, then reads each subsequent MBR more quickly because they know where to look for what they need.</p><p><strong>1. Headline and Investment Asks</strong> </p><p><strong>2. Delivery and Velocity</strong> </p><p><strong>3. Reliability and Operational Health</strong> </p><p><strong>4. Cost and Financial Impact</strong> </p><p><strong>5. Security, Compliance, and Risk</strong></p><p><strong>6. Next Month Priorities and Decisions Needed</strong></p><p>Each section follows a sub-structure: outcome statement, supporting metrics, narrative context, and any decision the section requires from leadership.</p><div><hr></div><h2>Section 1: Headline and Investment Asks</h2><p>The first 200 words of the MBR are the most important. If leadership stops reading after the first page, this section must contain the entire investment case.</p><p>The headline names the single most consequential outcome of the month in business terms. Not &#8220;we shipped the service mesh migration.&#8221; Instead: &#8220;MTTR for network-layer incidents dropped 60 percent following the service mesh migration, removing 3 hours of monthly engineer time previously spent on manual failover.&#8221;</p><p>Below the headline, three to five investment asks. Each ask is a single sentence stating the required decision, the timeframe, and the consequence of inaction. Examples:</p><ul><li><p>&#8220;Approve the additional senior platform engineer in Q3 planning. Without this hire, the multi-account migration delivery date moves from October to January.&#8221;</p></li><li><p>&#8220;Authorize $80,000 in annual spend for the observability platform upgrade. Current tooling cannot ingest the volume the new compute fleet generates, creating monitoring gaps in production.&#8221;</p></li><li><p>&#8220;Confirm the platform team owns the compliance infrastructure boundary. Ambiguity in the boundary ownership is delaying audit evidence collection by an average of 9 days per finding.&#8221;</p></li></ul><p>Each ask names the decision-maker implicitly: the VP, the CTO, the CFO. Leadership reads this section and knows exactly what is being asked of them.</p><div><hr></div><h2>Section 2: Delivery and Velocity</h2><p>This section reports what the platform shipped in business terms, not infrastructure terms.</p><p>Lead with two metrics: number of teams accelerated by platform work this month, and engineering days returned to product feature work. Both numbers must be defensible. If you cannot show the math, do not include the number.</p><p>A defensible delivery report looks like this:</p><p>&#8220;This month, the platform team delivered the standardized RDS provisioning pattern, which 8 product teams adopted. Each adopting team eliminated an average of 2.5 days previously spent on database setup, configuration review, and secrets management. Total engineering days returned to product roadmap work: 20 days.&#8221;</p><p>Below the metrics, list the three to five delivery items that produced the impact. Keep each item to one line. Do not list every Terraform module or ticket. The leader reading this is interpolating from your numbers rather than auditing your sprint board.</p><p>Close the section with one paragraph on velocity trends: is the platform team delivering more, less, or the same volume of impact compared to the prior quarter, and what is driving the trend?</p><div><hr></div><h2>Section 3: Reliability and Operational Health</h2><p>Reliability work is the most visible category of platform output but also the most prone to misreporting. The fix is to report both prevented and actual impact.</p><p>Lead with four metrics:</p><ul><li><p>Production incidents this month, with severity breakdown</p></li><li><p>Mean time to recovery, with comparison to the trailing 90-day average</p></li><li><p>Incidents prevented through platform-level mitigations (with the specific control and the incident path it blocked)</p></li><li><p>Service availability against SLO, by tier</p></li></ul><p>The number of prevented incidents is the most important. It is also the easiest to lose credibility if you overclaim. The discipline: only count an incident as prevented when there is a specific, named scenario that was caught in pre-production by a platform-owned control.</p><p>Example of a defensible prevented-incident entry: &#8220;On May 14, the IAM policy validator caught a privilege escalation path in a production deployment from the data-services team. Without the validator, the change would have granted s3:* on the customer-data bucket. Severity if shipped: P1.&#8221;</p><p>Close with a paragraph on the most consequential operational decision the team made this month, the reasoning, and the result.</p><div><hr></div><h2>Section 4: Cost and Financial Impact</h2><p>Cost is the section that most directly translates platform work into language the CFO recognizes. Most platform leaders under-report here because the wins feel speculative. They are not. They are calculable.</p><p>Three metrics anchor this section:</p><ul><li><p>Cloud spend this month vs prior month, with variance explanation</p></li><li><p>Cost optimizations realized this month, named individually with dollar impact</p></li><li><p>Forecast variance: Are you tracking to the annual cloud budget, and if not, what is driving the delta</p></li></ul><p>A complete cost optimization entry contains four pieces: what changed, how much it saved, when the saving begins, and who owns the next iteration. Example:</p><p>&#8220;Right-sized the analytics tier RDS instances from r6g.4xlarge to r6g.2xlarge after observing sustained CPU under 30 percent. Annualized savings: $58,000. Effective May 12. Next iteration: review the data-platform Aurora cluster sizing in Q3.&#8221;</p><p>Close the section with one observation on the cost trajectory and one cost decision that requires leadership input. Cost decisions almost always require leadership input because they involve trade-offs between performance, redundancy, and price.</p><div><hr></div><h2>Section 5: Security, Compliance, and Risk</h2><p>This section is where platform leaders systematically under-report and where leadership most often makes decisions on incomplete information.</p><p>Three categories belong here:</p><p><strong>Security posture changes.</strong> Specific controls added, removed, or modified. New attack paths closed. Vulnerabilities remediated with severity and exposure window.</p><p><strong>Compliance posture.</strong> Audit findings are open, in progress, and closed. Time-to-evidence by control. Compliance scope changes (new regions, new tenants, new frameworks).</p><p><strong>Risk register changes.</strong> Risks newly identified, risks materialized, risks closed. Each risk in the register has an owner, a likelihood, an impact, and a status.</p><p>The risk register is the most undervalued artifact in the MBR. Leadership cannot manage risk they do not see. A platform team that maintains a written risk register and updates it monthly converts itself from a delivery function to a governance function in the eyes of executive leadership.</p><p>A risk register entry looks like this: &#8220;Multi-region failover for the auth service is untested at production load. Likelihood: Medium. Impact: High (customer-facing outage if east-region fails). Owner: Platform team. Status: Tabletop scheduled for June 18. Decision needed: Confirm 4-hour engineer commitment from the SRE team for the test.&#8221;</p><div><hr></div><h2>Section 6: Next Month Priorities and Decisions Needed</h2><p>The final section closes the loop between this month&#8217;s results and next month&#8217;s commitments.</p><p>List the three to five priorities the platform team is committing to next month. Each priority states the outcome the team is targeting, the metric that will confirm success, and the date by which the work will be completed.</p><p>Below the priorities, list any decisions that leadership must make to enable the priorities. This is distinct from the investment asks at the top of the document. The questions at the top are organizational decisions about funding, headcount, and ownership. The decisions in this section are tactical: a specific approval, a specific resource, a specific stakeholder commitment.</p><p>Example: &#8220;To complete the SOC 2 evidence collection automation by June 28, the audit team must confirm the export format requirements by June 10. Without confirmation by June 10, the deliverable moves to July.&#8221;</p><p>This section makes the platform team&#8217;s dependencies on leadership explicit. It is also the section that most reliably produces follow-through, because leadership has already seen the math on what the dependency unblocks.</p><div><hr></div><h2>How to Run the MBR Cadence</h2><p>The document is half the work. The cadence around the document is the other half.</p><p><strong>Send the MBR three business days before the live review meeting.</strong> Leadership needs time to read, form questions, and check internal politics on any decisions you are asking for. A document delivered on the morning of the meeting forces leadership to react in the room, leading to worse decisions and eroding trust over time.</p><p><strong>The live meeting is 30 minutes, not 60.</strong> The document serves to convey information. The meeting exists to make decisions. If the meeting becomes a presentation of the document, the document has failed.</p><p><strong>Open the meeting with the investment asks, not the metrics.</strong> Leadership knows the metrics from the read-ahead. The meeting time is for resolving the asks. The first sentence in the room is &#8220;Section 1 lists three asks. Let&#8217;s start there.&#8221;</p><p><strong>Capture decisions in writing during the meeting.</strong> Send the decision summary to the same distribution list within 24 hours. The decision summary is what creates the audit trail and the accountability loop.</p><p>Run this cadence for three consecutive months before evaluating whether the format is working. </p><p>The first MBR will feel awkward. </p><p>The second will feel mechanical. </p><p>The third will start producing the conversations that change how the leadership funds the platform.</p><div><hr></div><h2>What to Measure to Evaluate the MBR Itself</h2><p>The MBR is a tool. Like any tool, it should be measured.</p><p><strong>Time-to-decision on investment asks.</strong> Track the date each ask is submitted and the date the decision is made. A healthy MBR cadence resolves requests within 15 business days. Asks that sit longer indicate either incomplete framing or gaps in escalation.</p><p><strong>Frequency of MBR-driven follow-up.</strong> Count the leadership questions, requests, or actions that originate from the MBR each month. A document that does not generate follow-up is being read but not acted on.</p><p><strong>Headcount and budget conversation outcomes.</strong> Track planning cycles. The MBR&#8217;s actual goal is to change the conversation in the next planning cycle. Within two cycles, platform headcount and budget discussions should reference MBR data points by name.</p><p>If none of these signals improve in the first six months, the format needs adjustment. Most often, the issue is one of three things: the headline is not specific enough, the investment asks are framed as wishes rather than decisions, or the risk register is missing.</p><div><hr></div><h2>When the MBR Becomes the Platform&#8217;s Operating Document</h2><p>The strongest signal that an MBR cadence is working is when leadership begins using it outside the meeting.</p><p>The CTO references the prior MBR in a board prep session. </p><p>The VP of Engineering pulls a metric from the MBR for a quarterly all-hands. </p><p>The CFO asks for the cost section formatted for a different audience. The MBR becomes the canonical record of what the platform team is, what it produces, and what it needs.</p><p>At that point, the document is no longer a reporting artifact. It is the operating document for the relationship between platform engineering and the rest of the executive team.</p><div><hr></div><h2>Upgrade If You Need Implementation, Not Just Ideas</h2><p>If you&#8217;re using these emails to guide real decisions on your platform, you&#8217;ll get more leverage from the paid version of The Cloud Playbook.</p><p>The free newsletter gives you patterns and language.</p><p>The paid newsletter turns those patterns into implementation kits you can ship inside a quarter:</p><ul><li><p>Concrete rollout plans (90&#8209;day roadmaps for each pattern)</p></li><li><p>Templates and checklists (policies, runbooks, tagging schemes, review checklists)</p></li><li><p>Real examples from high&#8209;stakes AWS environments (what we actually shipped and why)</p></li></ul><p>If the paid side doesn&#8217;t save you more than the subscription in <strong>one</strong> incident, audit cycle, or bad migration you avoid, you should cancel and keep the playbooks.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe&quot;,&quot;text&quot;:&quot;Upgrade to the Paid Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/subscribe"><span>Upgrade to the Paid Cloud Playbook</span></a></p><div><hr></div><h2><strong>That&#8217;s it for today!</strong></h2><p>Did you enjoy this newsletter issue?</p><p>Share with your friends, colleagues, and your favorite social media platform.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Cloud Playbook</span></a></p><p><strong>Until next week &#8212; Amrut</strong></p><div><hr></div><h2><strong>Get in touch</strong></h2><p>You can find me on <a href="https://www.linkedin.com/in/patilamrut/">LinkedIn</a> or <a href="https://twitter.com/realamrutpatil">X</a>.</p><p>If you would like to request a topic to read, please feel free to contact me directly via LinkedIn or X.</p>]]></content:encoded></item><item><title><![CDATA[TCP# 124: The 6 golden paths worth building first, and how to sequence them.]]></title><description><![CDATA[Scored by frequency, friction, risk, and reach. With backlog template and adoption metrics.]]></description><link>https://www.thecloudplaybook.com/p/first-6-golden-paths-aws-platform-teams</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/first-6-golden-paths-aws-platform-teams</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Sun, 24 May 2026 15:01:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!XSxO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a2a1509-c22d-4960-8e5a-4aa4129cda82_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Platform teams rarely have a golden path shortage. They have a sequencing problem.</p><p>The backlog of potential paths is long: containerized deployments, serverless functions, database provisioning, secrets management, observability setup, CI/CD templates. Every senior engineer has a candidate. Every app team has a complaint about the current manual process.</p><p>The question is not what to build. It is what to build first.</p><p>Wrong sequencing produces low adoption. A golden path for Kubernetes service mesh configuration is technically impressive and practically irrelevant if most teams are still deploying Lambda functions manually. App teams ignore paths that don&#8217;t match their daily work. The platform team loses credibility before the program gets traction.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XSxO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a2a1509-c22d-4960-8e5a-4aa4129cda82_1254x1254.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XSxO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a2a1509-c22d-4960-8e5a-4aa4129cda82_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!XSxO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a2a1509-c22d-4960-8e5a-4aa4129cda82_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!XSxO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a2a1509-c22d-4960-8e5a-4aa4129cda82_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!XSxO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a2a1509-c22d-4960-8e5a-4aa4129cda82_1254x1254.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XSxO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a2a1509-c22d-4960-8e5a-4aa4129cda82_1254x1254.png" width="1254" height="1254" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7a2a1509-c22d-4960-8e5a-4aa4129cda82_1254x1254.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1254,&quot;width&quot;:1254,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1492693,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/197148237?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a2a1509-c22d-4960-8e5a-4aa4129cda82_1254x1254.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!XSxO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a2a1509-c22d-4960-8e5a-4aa4129cda82_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!XSxO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a2a1509-c22d-4960-8e5a-4aa4129cda82_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!XSxO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a2a1509-c22d-4960-8e5a-4aa4129cda82_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!XSxO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a2a1509-c22d-4960-8e5a-4aa4129cda82_1254x1254.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Why the Sequence of Golden Paths Determines Platform Trust</h2><p>The first three paths a platform team ships set the tone for everything that follows.</p><p>If the first paths are fast, correct, and clearly easier than the alternative, app teams adopt them. They stop routing around the platform. They start asking for the next path instead of building their own. The platform team earns the right to set standards for more complex decisions.</p><p>If the first paths are slow to produce, hard to use, or misaligned with what teams actually need in their daily work, adoption fails. Platform engineering becomes the team that built tools nobody uses. Credibility takes quarters to rebuild.</p><p>The first six paths are not a starting point. They are the foundation of platform trust.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Cloud Playbook! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>The Failure Mode: Building for Interest Instead of Impact</h2><p>Most platform teams prioritize golden paths based on what is technically interesting to build, not on what creates the most daily impact for app teams.</p><p>They build a Kubernetes deployment path because the platform engineers are skilled in Kubernetes. They build an advanced multi-region failover path because it is architecturally sophisticated. They build a developer portal integration because it makes a good conference talk.</p><p>Meanwhile, every app team is manually configuring Lambda functions with inconsistent IAM roles, provisioning RDS instances without encryption enforcement, and setting up CloudWatch alarms from scratch for every new service. These are the gaps producing incidents, cost anomalies, and audit findings right now.</p><p>A golden path program that prioritizes interesting work over daily friction earns goodwill from platform engineers and indifference from app teams.</p><div><hr></div><h2>A Scoring Model for Golden Path Prioritization</h2><p>Score each candidate path across four dimensions before committing the build investment.</p><ul><li><p><strong>Frequency.</strong> How many engineers would use this path in a given month? A path used daily by every team outranks one used once per quarter by a single team. Count the actual provisioning events, not the theoretical ones.</p></li><li><p><strong>Friction.</strong> How much time does the current manual approach cost? A path that eliminates 90 minutes of manual configuration per deployment creates more measurable value than a path that eliminates 10 minutes.</p></li><li><p><strong>Risk.</strong> What goes wrong when teams build without this path? Score on two sub-dimensions: incident risk (how often has the manual approach produced a production failure?) and audit risk (how often has the manual approach produced a compliance gap?).</p></li><li><p><strong>Reach.</strong> How many existing services would benefit from retroactive adoption? A path that applies to 40 current services without re-provisioning creates value beyond new deployments.</p></li></ul><p>The six paths below score highest across these four dimensions in most AWS SaaS environments.</p><div><hr></div><h2>The Six Paths, Scored and Sequenced</h2><p><strong>Path 1: Containerized Service Deployment (ECS on Fargate)</strong></p><p>The highest-frequency deployment pattern for most SaaS platform teams. Manual deployment requires engineers to configure the task definition, ECS service, target group, log group, IAM execution role, and cost allocation tags independently and consistently.</p><ul><li><p>Owner: Platform Engineering, compute chapter</p></li><li><p>What the module covers: Task definition template, ECS service configuration, ALB target group, CloudWatch log group with standard retention, IAM execution role with least-privilege policy, required cost tags</p></li><li><p>Adoption metric: Percentage of new ECS services deployed via the module in the trailing 30 days</p></li><li><p>Rollout sequence: Release to one app team as a pilot, collect feedback on friction points, iterate, then open to all teams</p></li></ul><p><strong>Path 2: Serverless Function Deployment (Lambda)</strong></p><p>Lambda is the second-highest-frequency deployment pattern and the one most likely to have inconsistent IAM scopes and missing observability in existing deployments.</p><ul><li><p>Owner: Platform Engineering, serverless chapter</p></li><li><p>What the module covers: Function configuration with runtime defaults, IAM execution role scoped to required permissions, CloudWatch log group with retention policy, error rate alarm with defined threshold, required cost and service tags</p></li><li><p>Adoption metric: Percentage of new Lambda functions with IAM roles sourced from the module&#8217;s policy template</p></li><li><p>Rollout sequence: Pilot with the team that has the most Lambda functions in production, prioritize retroactive adoption to close existing IAM and observability gaps</p></li></ul><p><strong>Path 3: Managed Relational Database Provisioning (RDS)</strong></p><p>Database provisioning has the highest incident and audit risk among self-service infrastructure decisions. Encryption at rest, backup retention, deletion protection, subnet group placement, and parameter group selection are regularly misconfigured by teams building in the AWS console or from memory.</p><ul><li><p>Owner: Platform Engineering, data chapter</p></li><li><p>What the module covers: Instance class guardrails by environment tier, encryption at rest enforced, automated backup enabled with defined retention period, deletion protection enabled in production, subnet group assignment to private subnets, parameter group set to approved baseline, required tags</p></li><li><p>Adoption metric: Percentage of RDS instances in production accounts that were provisioned via the module</p></li><li><p>Rollout sequence: Release to production after extensive testing in staging, prioritize retroactive compliance for existing instances</p></li></ul><p><strong>Path 4: Cost-Tagged Infrastructure Module</strong></p><p>Every resource type in your environment should apply a consistent set of cost allocation tags. The alternative is a quarterly archaeology project to attribute spend. This path is not a deployment module for a specific resource. It is a tagging standard embedded in every other module.</p><ul><li><p>Owner: Platform Engineering, FinOps chapter</p></li><li><p>What the module covers: Required tag definitions (team, product, environment, cost-center, service-owner), tag validation logic, tag policy enforcement via AWS Config or Terraform plan validation, tag propagation to child resources</p></li><li><p>Adoption metric: Percentage of resources in production accounts that carry all required cost allocation tags</p></li><li><p>Rollout sequence: Release as a dependency of Paths 1, 2, and 3 &#8212; embedded, not optional</p></li></ul><p><strong>Path 5: Observability Baseline for New Services</strong></p><p>Every new service needs CloudWatch metrics, structured log groups, distributed tracing, and at least two alarms: one on error rate and one on latency. Teams building without this path produce services that are invisible until they fail.</p><ul><li><p>Owner: Platform Engineering, observability chapter</p></li><li><p>What the module covers: CloudWatch metric namespace for the service, log group with structured format and retention policy, X-Ray tracing enabled, error rate alarm, p99 latency alarm, dashboard template for the service&#8217;s health view</p></li><li><p>Adoption metric: Percentage of services in production with all five observability components present</p></li><li><p>Rollout sequence: Release alongside Path 1 and Path 2 so new service deployments include observability by default</p></li></ul><p><strong>Path 6: Secrets Management Pattern (AWS Secrets Manager)</strong></p><p>Hardcoded credentials and environment variables containing secrets are among the most common audit findings on cloud platforms. This path eliminates the decision: secrets go in Secrets Manager, with rotation enabled, and the application accesses them via a scoped IAM policy.</p><ul><li><p>Owner: Platform Engineering, security chapter</p></li><li><p>What the module covers: Secrets Manager secret provisioning with rotation enabled, rotation Lambda for supported secret types, IAM policy granting read access scoped to the specific secret, resource policy blocking cross-account access by default</p></li><li><p>Adoption metric: Percentage of application secrets accessed via Secrets Manager vs. environment variables or parameter store without rotation</p></li><li><p>Rollout sequence: Pilot with one team migrating hardcoded credentials, document the migration steps, then open as the default secret provisioning path</p></li></ul><div><hr></div><h2>Artifact in This Issue</h2><p>The artifact is the Golden Path Backlog Template: a structured planning table for sequencing your golden path program.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://drive.google.com/file/d/1Uc8PteJrY2IewUiNJtgr-E4CKLdc3knc/view?usp=sharing&quot;,&quot;text&quot;:&quot;Download&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://drive.google.com/file/d/1Uc8PteJrY2IewUiNJtgr-E4CKLdc3knc/view?usp=sharing"><span>Download</span></a></p><p>Each row in the template contains:</p><ul><li><p>Path name and type (deployment, provisioning, configuration, observability)</p></li><li><p>Problem solved: the specific manual process this path replaces</p></li><li><p>Frequency score: monthly provisioning events across the engineering org</p></li><li><p>Friction score: estimated hours of manual work eliminated per use</p></li><li><p>Risk score: incident and audit risk rating from the manual approach</p></li><li><p>Reach score: number of existing services that can retroactively adopt the path</p></li><li><p>Total priority score and recommended sequencing position</p></li><li><p>Owner team and implementation scope summary</p></li><li><p>Adoption metric definition</p></li><li><p>Rollout sequence steps (pilot, feedback, iterate, release)</p></li></ul><p>Use this template in your next platform engineering planning session to build your own prioritized golden path backlog. Score each candidate path against your actual provisioning data, not estimates. The team that deploys 20 Lambda functions per month has a higher friction score for Path 2 than the team that deploys two.</p><div><hr></div><h2>What to Measure and When to Review</h2><ul><li><p><strong>Adoption rate per path.</strong> Percentage of new deployments using the golden path module in the trailing 30 days. Track per path, not as a single aggregate. Paths with adoption below 60 percent within 90 days of release need a friction audit.</p></li><li><p><strong>Drift rate.</strong> Percentage of existing services that were compliant with the path standard at release but have since drifted. High drift indicates the path lacks enforcement mechanisms. It is documentation, not a system.</p></li><li><p><strong>Time saved per deployment.</strong> Estimate the manual configuration time eliminated by each path use. Accumulate monthly. This number converts golden path investment into engineering hours returned to product work, which is the right currency for the conversation with the VP of Engineering.</p></li></ul><p>Review path adoption metrics monthly in the platform team&#8217;s operating review. Review the prioritized backlog quarterly to add candidates, retire unused paths, and promote complex paths to simpler defaults as adoption matures.</p><div><hr></div><h2>What the First Six Paths Actually Build</h2><p>The six paths above are not primarily about technical consistency, though they produce it.</p><p>They are about demonstrating that the platform team solves real problems fast, makes the right choice the easy choice, and delivers value that app teams can measure in time and incidents.</p><p>When the first six paths land well, app teams stop routing around the platform. They start requesting paths 7 through 12. The platform team stops defending its investment and starts negotiating for the capacity to expand it.</p><p>The path program that scales is not the one with the most sophisticated architecture. It is the one that earned trust early by solving the problems teams had every day.</p><div><hr></div><p>Use the backlog template as your agenda for next quarter&#8217;s planning. Score your top ten candidate paths before the session so the team is debating real data, not intuition. Forward this issue to the senior engineers and EMs who own the highest-friction provisioning workflows on their teams. They are the ones who know where the manual time is going.</p><div><hr></div><h2><strong>That&#8217;s it for today!</strong></h2><p>Did you enjoy this newsletter issue?</p><p>Share with your friends, colleagues, and your favorite social media platform.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Cloud Playbook</span></a></p><p><strong>Until next week &#8212; Amrut</strong></p><div><hr></div><h2><strong>Get in touch</strong></h2><p>You can find me on <a href="https://www.linkedin.com/in/patilamrut/">LinkedIn</a> or <a href="https://twitter.com/realamrutpatil">X</a>.</p><p>If you would like to request a topic to read, please feel free to contact me directly via LinkedIn or X.</p>]]></content:encoded></item><item><title><![CDATA[TCP #123: Golden paths fail when they require engineers to choose them]]></title><description><![CDATA[The difference between documentation and a system, and why one scales while the other doesn't.]]></description><link>https://www.thecloudplaybook.com/p/golden-paths-documentation-not-systems</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/golden-paths-documentation-not-systems</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Sun, 17 May 2026 14:30:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4b58!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd43fd147-24b1-476b-a763-95ea6b1e26ee_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The platform team publishes the golden path. </p><p>A Confluence page. A README. A wiki entry explaining the approved way to deploy a new service, provision a database, or configure observability.</p><p>Adoption is strong in the first two weeks. Engineers read it during onboarding. A few teams follow it on their next service.</p><p>Six months later, the platform team runs an audit. Half the services in production do not comply with the standard. Some never did. Engineers who read the doc once are deploying from memory. Teams under a deadline skipped the path entirely and built what they already knew.</p><p>The platform team calls it an adoption problem. It is not. It is a design problem.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4b58!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd43fd147-24b1-476b-a763-95ea6b1e26ee_1254x1254.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4b58!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd43fd147-24b1-476b-a763-95ea6b1e26ee_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!4b58!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd43fd147-24b1-476b-a763-95ea6b1e26ee_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!4b58!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd43fd147-24b1-476b-a763-95ea6b1e26ee_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!4b58!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd43fd147-24b1-476b-a763-95ea6b1e26ee_1254x1254.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4b58!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd43fd147-24b1-476b-a763-95ea6b1e26ee_1254x1254.png" width="1254" height="1254" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d43fd147-24b1-476b-a763-95ea6b1e26ee_1254x1254.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1254,&quot;width&quot;:1254,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1467493,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/197147610?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd43fd147-24b1-476b-a763-95ea6b1e26ee_1254x1254.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4b58!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd43fd147-24b1-476b-a763-95ea6b1e26ee_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!4b58!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd43fd147-24b1-476b-a763-95ea6b1e26ee_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!4b58!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd43fd147-24b1-476b-a763-95ea6b1e26ee_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!4b58!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd43fd147-24b1-476b-a763-95ea6b1e26ee_1254x1254.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The Business Case for Getting Golden Paths Right</h2><p>Every decision an engineer makes without a golden path is a chance to deviate from the standard.</p><p>Manual decisions slow delivery. An engineer building a new Lambda function from scratch spends two hours figuring out the right IAM role scope, log retention policy, and error alerting configuration. A golden path that wraps this into a validated module turns that into a 15-minute task.</p><p>Manual decisions create inconsistency. Inconsistency compounds. Twelve teams deploying services twelve different ways means twelve different log formats, twelve different tagging conventions, and twelve different security postures to audit against.</p><p>Inconsistency produces three outcomes that appear on the platform team&#8217;s desk: incidents from configurations that deviated from a tested baseline, AWS cost anomalies from resource patterns that bypassed the approved cost controls, and audit findings from services that were never reviewed against the security standard.</p><p>The golden path does not eliminate these problems by existing. It eliminates them by being used.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Cloud Playbook is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Why Documentation Fails Under Pressure</h2><p>A Confluence page requires the engineer to do five things before the standard applies: know the page exists, find it, read it, remember the relevant parts, and choose to follow it when under deadline pressure with a closing release window.</p><p>That is five failure points. Under a deadline, discipline competes with speed. Speed wins.</p><p>Documentation-based golden paths suffer from three structural weaknesses that no amount of engineering culture can fix.</p><ul><li><p><strong>They require active adoption</strong></p><p>The engineer must choose the golden path each time. In low-pressure conditions, most do. Under a production incident, a Friday release, or a late-quarter launch push, most don&#8217;t. The path that requires a decision is the path that gets bypassed.</p></li><li><p><strong>They drift from the standard without detection</strong></p><p>A Confluence page can describe the right pattern for provisioning a database. It cannot prevent an engineer from provisioning a different pattern. The documentation and the deployed state diverge silently. The platform team discovers the gap at the next audit.</p></li><li><p><strong>They have no adoption signal</strong></p><p>You cannot tell from a Confluence page how many teams followed it, which teams ignored it, or when the last adoption happened. Without a signal, the platform team cannot distinguish between &#8220;the path is working&#8221; and &#8220;the path was abandoned months ago.&#8221;</p></li></ul><div><hr></div><h2>Four Axes for Evaluating Whether Your Golden Path Is a System</h2><p>A golden path becomes a system when it reduces the cost of the right choice to below that of any alternative. Evaluate your existing golden paths against four dimensions:</p><ul><li><p><strong>Friction.</strong> Does using the golden path take less time than not using it? If the golden path requires more steps than the manual alternative, engineers will take the shorter route. A Terraform module that provisions a correctly configured RDS instance in five minutes beats a doc that explains how to configure one correctly in forty-five.</p></li><li><p><strong>Decision elimination.</strong> Does the golden path remove choices, or does it document the right choice and leave the decision to the engineer? The distinction matters. A module that enforces the encryption setting eliminates the need for a decision. A doc that says &#8220;enable encryption&#8221; documents it. One applies the standard automatically. The other relies on the engineer remembering to do so.</p></li><li><p><strong>Drift resistance.</strong> Can the golden path drift from the standard over time, or does it enforce the standard automatically? A Terraform module version-pinned to a tested configuration drifts when someone modifies it. A policy-enforced guardrail that rejects non-compliant resources does not drift. Drift resistance is a function of the enforcement mechanism, not documentation quality.</p></li><li><p><strong>Adoption measurement.</strong> Do you know who is using the golden path, how often, and which services have adopted it? An internal module registry that tracks download counts, a pipeline template that logs invocations, or a tag applied automatically on module use all produce adoption signals. A Confluence page produces none.</p></li></ul><p>A golden path that scores well across all four dimensions does not require marketing effort. It gets used because it is the lowest-friction option available.</p><div><hr></div><h2>What Platform Teams That Get This Right Actually Build</h2><p>When golden paths are systems, three things change.</p><p>First, adoption is automatic. Engineers do not choose the golden path. They use it because it is the default, the fastest option, and the only path that meets compliance requirements without additional configuration.</p><p>Second, drift detection becomes operational. When services must pass through a validated module or pipeline template, deviations appear in the deployment log rather than in the audit. The platform team catches the gap before the auditor does.</p><p>Third, investment becomes measurable. Adoption metrics show which paths are in active use, which teams are using them, and what the alternative deployment time would have been. The platform team can quantify the hours saved, the incidents prevented, and the reduced audit exposure. That evidence changes the conversation with the CTO from &#8220;justify the platform investment&#8221; to &#8220;where should we build the next path.&#8221;</p><div><hr></div><p><em>On Wednesday, paid subscribers get the first 6 golden paths every AWS platform team should build, including owner, implementation scope, adoption metric, and rollout sequence for each. The issue prioritizes paths by the combination of daily usage frequency, manual time cost, and compliance risk if skipped.</em></p><h2>Upgrade If You Need Implementation, Not Just Ideas</h2><p>If you&#8217;re using these emails to guide real decisions on your platform, you&#8217;ll get more leverage from the paid version of The Cloud Playbook.</p><p>The free newsletter gives you patterns and language.</p><p>The paid newsletter turns those patterns into implementation kits you can ship inside a quarter:</p><ul><li><p>Concrete rollout plans (90&#8209;day roadmaps for each pattern)</p></li><li><p>Templates and checklists (policies, runbooks, tagging schemes, review checklists)</p></li><li><p>Real examples from high&#8209;stakes AWS environments (what we actually shipped and why)</p></li></ul><p>If the paid side doesn&#8217;t save you more than the subscription in <strong>one</strong> incident, audit cycle, or bad migration you avoid, you should cancel and keep the playbooks.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe&quot;,&quot;text&quot;:&quot;Upgrade to the Paid Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thecloudplaybook.com/subscribe"><span>Upgrade to the Paid Cloud Playbook</span></a></p><div><hr></div><h2><strong>That&#8217;s it for today!</strong></h2><p>Did you enjoy this newsletter issue?</p><p>Share with your friends, colleagues, and your favorite social media platform.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Cloud Playbook</span></a></p><p><strong>Until next week &#8212; Amrut</strong></p><div><hr></div><h2><strong>Get in touch</strong></h2><p>You can find me on <a href="https://www.linkedin.com/in/patilamrut/">LinkedIn</a> or <a href="https://twitter.com/realamrutpatil">X</a>.</p><p>If you would like to request a topic to read, please feel free to contact me directly via LinkedIn or X.</p>]]></content:encoded></item><item><title><![CDATA[TCP #122: Your approval process needs a classification model, not just a queue]]></title><description><![CDATA[Three-tier framework, change matrix, and checklist for what platform must own, review, or release.]]></description><link>https://www.thecloudplaybook.com/p/platform-approval-model-change-classification</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/platform-approval-model-change-classification</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Wed, 13 May 2026 13:02:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5NND!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa369f3a6-2ba5-4b0d-9ec2-dbb743829cd2_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When platform teams grow their approval scope without a classification model, two things happen simultaneously.</p><p>App teams wait for approvals they do not actually need. And high-risk infrastructure changes move through the same queue as routine additions, reviewed at the same depth and with the same SLA.</p><p>The result is that the platform is slow, and the risks it was built to catch are still getting through.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5NND!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa369f3a6-2ba5-4b0d-9ec2-dbb743829cd2_1254x1254.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5NND!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa369f3a6-2ba5-4b0d-9ec2-dbb743829cd2_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!5NND!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa369f3a6-2ba5-4b0d-9ec2-dbb743829cd2_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!5NND!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa369f3a6-2ba5-4b0d-9ec2-dbb743829cd2_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!5NND!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa369f3a6-2ba5-4b0d-9ec2-dbb743829cd2_1254x1254.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5NND!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa369f3a6-2ba5-4b0d-9ec2-dbb743829cd2_1254x1254.png" width="1254" height="1254" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a369f3a6-2ba5-4b0d-9ec2-dbb743829cd2_1254x1254.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1254,&quot;width&quot;:1254,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1481193,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/197146000?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa369f3a6-2ba5-4b0d-9ec2-dbb743829cd2_1254x1254.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5NND!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa369f3a6-2ba5-4b0d-9ec2-dbb743829cd2_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!5NND!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa369f3a6-2ba5-4b0d-9ec2-dbb743829cd2_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!5NND!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa369f3a6-2ba5-4b0d-9ec2-dbb743829cd2_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!5NND!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa369f3a6-2ba5-4b0d-9ec2-dbb743829cd2_1254x1254.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Why Classification Determines Platform Scale</h2><p>Every change without a classification model lands in the same queue. The platform engineer reviews it, asks the same baseline questions, and applies the same scrutiny regardless of actual risk.</p><p>This works for ten engineers. At forty, the queue is full before Tuesday morning. App teams submit on Monday and receive responses on Thursday. They learn to route around the process. They provision what they need outside the intake channel, outside tagging standards, outside the platform&#8217;s field of view.</p><p>The CTO sees slow delivery and a platform team that cannot explain its own backlog. The CFO sees an AWS bill with no clear attribution. The platform team answers both questions without controlling either outcome.</p><p>The root cause is not throughput. It is a classification. The platform is reviewing decisions that should never require review while inadequately scrutinizing those that genuinely do.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Cloud Playbook! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>The Two Failures That Keep Teams Here</h2><div class="paywall-jump" data-component-name="PaywallToDOM"></div><p>The first failure is binary thinking: every change either requires platform approval or it does not. Binary classification turns every edge case into a judgment call. Reviewers decide inconsistently. App teams learn which answer to expect and route requests accordingly. The approval process becomes a negotiation rather than a standard.</p><p>The second failure is calibrating risk on surface features rather than consequences. A complex Terraform file looks risky. A tag update looks trivial. But a missing cost allocation tag on a production RDS instance can result in months of misattributed spend. A well-validated Terraform module for a standard ECS service requires no review.</p><p>Risk lives in three variables: blast radius, reversibility, and compliance scope. Not in file complexity or line count. A classification model built on those three variables outperforms any heuristic based on surface appearance.</p><div><hr></div><h2>A Three-Tier Change Classification Model</h2><p>Tier classification sorts every change into one of three buckets based on risk profile, not technical complexity.</p><p><strong>Tier 1: Self-Service</strong></p><p>Changes that fall within established standards affect only the requesting team&#8217;s scope and are easily reversible if incorrect. No platform review required. Examples:</p><ul><li><p>Scaling compute within an approved instance family</p></li><li><p>Adding resources using approved Terraform modules with required cost allocation tags applied</p></li><li><p>Modifying application configuration for services already in scope</p></li><li><p>Updating routing rules within team-owned load balancers</p></li></ul><p>App teams document Tier 1 changes in their own change log. Platform audits a sample monthly, not every instance.</p><p><strong>Tier 2: Platform Review Required</strong></p><p>Changes that introduce new patterns, cross team boundaries, affect shared infrastructure, or touch controls within compliance scope. Examples:</p><ul><li><p>New VPC configurations or subnet additions</p></li><li><p>IAM roles with cross-account trust or elevated permissions</p></li><li><p>S3 bucket creation in production accounts</p></li><li><p>RDS instance provisioning above the defined size thresholds</p></li><li><p>Security group modifications affecting shared services</p></li><li><p>Any change to networking or compute that modifies a control in scope for SOC 2, FedRAMP, HIPAA, or ISO 27001</p></li></ul><p>Platform reviews Tier 2 changes within a defined SLA: one business day for standard configurations, three business days for requests requiring compliance assessment. The sensitive-change checklist governs what reviewers verify before approving.</p><p><strong>Tier 3: Platform-Owned Changes</strong></p><p>Changes that must never leave the platform's hands. These are decisions with a wide blast radius, direct compliance control, ownership, or irreversible consequences if wrong. Examples:</p><ul><li><p>Changes to account-level SCPs</p></li><li><p>Modifications to centralized logging or audit trail infrastructure</p></li><li><p>VPC peering, Transit Gateway, or Direct Connect configuration</p></li><li><p>Encryption key management and rotation policy</p></li><li><p>Organization-level IAM or identity federation configuration</p></li><li><p>Backup and disaster recovery configuration for shared infrastructure</p></li></ul><p>App teams do not submit Tier 3 items as requests. They describe the business need. Platform engineers own the implementation.</p><p>The key question for classifying any change between Tier 1 and Tier 2: if this configuration is wrong, how long will it take for someone to detect the impact, and how hard will it be to reverse? That question drives tier placement more reliably than any checklist.</p><div><hr></div><h2>Building the Classification System at Your Platform</h2><p>Follow this sequence to implement change classification:</p><p><strong>1. List your accountability scope explicitly.</strong> What does your platform team actually own? Reliability SLAs, cloud cost reporting, compliance posture, networking, identity, observability? Write it as a named list. This list becomes the anchor for the entire model.</p><p><strong>2. Map resource types to accountability areas.</strong> For each area, identify which AWS resource types and configuration choices directly affect the outcome. Cost accountability maps to compute sizing, storage provisioning, and cost allocation tagging. Compliance scope maps to encryption settings, access control, logging destinations, and network exposure.</p><p><strong>3. Score each resource type for blast radius and reversibility.</strong> Blast radius: if this resource is misconfigured, what breaks? Reversibility: if the error is caught after deployment, how quickly can it be corrected without impacting production? High blast radius plus low reversibility equals Tier 3. Low blast radius plus high reversibility equals Tier 1.</p><p><strong>4. Define the self-service boundary explicitly.</strong> Publish the list of resources and configurations that app teams can provision without review. Be specific. &#8220;EC2 instances using approved AMIs within approved instance families, tagged with required cost allocation tags, deployed via approved Terraform modules, into pre-approved subnets&#8221; is a self-service definition. &#8220;Standard EC2 instances&#8221; is not.</p><p><strong>5. Build the sensitive-change checklist.</strong> For each Tier 2 resource category, define the specific items a reviewer verifies before approving. The checklist for IAM role creation differs from the checklist for RDS provisioning. Keep each checklist under 10 items. More than that signals the category needs further decomposition.</p><p><strong>6. Name the escalation path.</strong> When a reviewer is uncertain about tier placement for an ambiguous request, who makes the final call? Name that person and document the process before the first edge case arrives.</p><div><hr></div><h2>Artifact in This Issue</h2><p>The artifact is the <strong>Platform Change Classification Matrix</strong>: a structured reference table for categorizing infrastructure changes by tier, risk profile, and review requirements.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://drive.google.com/file/d/1SHX2S_9M83yDxue3TDmFWfwM6NzmUZ4F/view?usp=sharing&quot;,&quot;text&quot;:&quot;Download&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://drive.google.com/file/d/1SHX2S_9M83yDxue3TDmFWfwM6NzmUZ4F/view?usp=sharing"><span>Download</span></a></p><p>The matrix contains:</p><ul><li><p>Resource type or change category (IAM role creation, S3 bucket provisioning, SCP modification, and 20 additional common resource types)</p></li><li><p>Risk indicators for each: blast radius score, reversibility score, compliance scope flag, shared infrastructure flag</p></li><li><p>Tier assignment: Self-Service, Platform Review, or Platform-Owned</p></li><li><p>Review the owner and review the SLA for every Tier 2 entry</p></li><li><p>Sensitive-change checklist for each Tier 2 category, listing the specific items a reviewer verifies before approving</p></li></ul><p>The matrix is structured as a flat reference table. Sort by tier to give reviewers a quick-reference view. Sort by resource type to give app teams a self-service lookup.</p><p>Use this matrix as the starting point for your next platform ops review session. Walk through the resource types in your environment, assign each to a tier, and resolve disagreements using the blast radius and reversibility scoring. Publish the completed version in your internal wiki and reference it as the first step in your intake form before any request enters the queue.</p><div><hr></div><h2>What to Measure and When to Review</h2><p><strong>Classification distribution.</strong> Track the percentage of intake requests landing in each tier each week. A healthy distribution: 60 to 70 percent Tier 1, 25 to 35 percent Tier 2, 5 to 10 percent Tier 3. If Tier 1 is below 50 percent, the self-service boundary is too narrow. If Tier 2 exceeds 50 percent, audit whether reviewers are escalating ambiguous requests rather than approving or rejecting them.</p><p><strong>Review cycle time for Tier 2.</strong> Mean time from submission to approval decision. Set a target SLA and track it weekly. If cycle time consistently exceeds the SLA, investigate whether the sensitive-change checklist is calibrated correctly or whether submitters are arriving with incomplete context.</p><p>Review the matrix quarterly. As your approved Terraform module library grows and app teams become familiar with standards, Tier 2 items should migrate to Tier 1. The model should become less restrictive over time as platform maturity increases.</p><div><hr></div><h2>What Changes When Classification Is Right</h2><p>When the classification model is working, the platform team is no longer the one slowing delivery. App teams run self-service for their own decisions. Reviewers focus on the changes that actually need review. Platform-owned decisions stay in the platform's hands.</p><p>Platform engineers stop triaging an undifferentiated backlog. They start doing architecture review work worth doing.</p><p>The compliance evidence improves. Tier 3 changes leave a clean ownership trail. Tier 2 approvals are documented against a checklist. Tier 1 changes are auditable via sampling rather than exhaustive review.</p><p>The platform team builds a reputation for predictable, fast responses. That reputation is more durable than any SLA commitment made without a supporting model.</p><div><hr></div><p>Use this matrix as the agenda for your next platform ops session. </p><p>Walk your team through the resource types in your environment and assign tiers. Share the completed matrix with the EMs and senior engineers who submit intake requests. </p><p>Run a calibration pass after the first 30 days: sort the previous month&#8217;s requests by tier and verify the classifications held. </p><p>If you manage a platform for multiple product teams, forward this to the EM or senior engineer who will be the primary intake submitter on each team.</p><div><hr></div><h2><strong>That&#8217;s it for today!</strong></h2><p>Did you enjoy this newsletter issue?</p><p>Share with your friends, colleagues, and your favorite social media platform.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Cloud Playbook</span></a></p><p><strong>Until next week &#8212; Amrut</strong></p><div><hr></div><h2><strong>Get in touch</strong></h2><p>You can find me on <a href="https://www.linkedin.com/in/patilamrut/">LinkedIn</a> or <a href="https://twitter.com/realamrutpatil">X</a>.</p><p>If you would like to request a topic to read, please feel free to contact me directly via LinkedIn or X.</p>]]></content:encoded></item><item><title><![CDATA[TCP #121: Accountability Without Authority Is How Platform Teams Fail]]></title><description><![CDATA[When platform is judged on reliability, cost, and compliance without approval rights over infrastructure, failure is structural, not personal.]]></description><link>https://www.thecloudplaybook.com/p/platform-accountability-without-authority-failure</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/platform-accountability-without-authority-failure</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Sun, 10 May 2026 14:30:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_OTe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d94943-f544-4e00-9422-58a937be63ad_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The most common dysfunction I see in platform engineering is not technical.</p><p>It is organizational.</p><p>The platform team is accountable for reliability, cost, and compliance readiness. Simultaneously, app teams provision their own infrastructure, configure their own environments, and make their own architecture choices. The platform has no approval rights over those decisions.</p><p>When something breaks, the platform team explains the incident. When the AWS bill is too high, the platform team presents the cost review. When the auditor finds a misconfigured S3 bucket provisioned by an app team, the platform team answers for it.</p><p>This is not a people problem. It is a structural mismatch: accountability without authority.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_OTe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d94943-f544-4e00-9422-58a937be63ad_1254x1254.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_OTe!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d94943-f544-4e00-9422-58a937be63ad_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!_OTe!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d94943-f544-4e00-9422-58a937be63ad_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!_OTe!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d94943-f544-4e00-9422-58a937be63ad_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!_OTe!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d94943-f544-4e00-9422-58a937be63ad_1254x1254.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_OTe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d94943-f544-4e00-9422-58a937be63ad_1254x1254.png" width="1254" height="1254" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0d94943-f544-4e00-9422-58a937be63ad_1254x1254.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1254,&quot;width&quot;:1254,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1487684,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/196371584?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d94943-f544-4e00-9422-58a937be63ad_1254x1254.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_OTe!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d94943-f544-4e00-9422-58a937be63ad_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!_OTe!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d94943-f544-4e00-9422-58a937be63ad_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!_OTe!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d94943-f544-4e00-9422-58a937be63ad_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!_OTe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d94943-f544-4e00-9422-58a937be63ad_1254x1254.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The Business Cost of Structural Mismatch</h2><p>Structural mismatch in platform engineering manifests in three ways, all of which are costly.</p><ul><li><p><strong>Reliability incidents you cannot prevent</strong></p></li></ul><p>App teams deploy changes outside the platform&#8217;s change management process. Those changes introduce instability. Platform is on call for the fallout. The platform team cannot stop the root cause; they can only respond to it after the fact.</p><ul><li><p><strong>Cloud cost variance you cannot explain</strong></p></li></ul><p>App teams create resources without tagging standards. The platform team reports total cloud spend to the CTO, but they cannot attribute it to teams, products, or tenants with confidence. Cost reviews become estimates. Budget conversations become defensive.</p><ul><li><p><strong>Audit findings you cannot remediate</strong></p></li></ul><p>A control requires that all S3 buckets be encrypted. An app team creates a bucket without it. The auditor finds it. The platform team owns the control. But they did not own the bucket.</p><p>Each of these is a variation of the same root problem. The platform owns the outcome without owning the inputs.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Cloud Playbook is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>How Smart Teams End Up Here</h2><p>This structure does not happen by accident. It usually develops through a reasonable sequence of decisions.</p><p>The company starts small. Developers provision their own infrastructure. It works. Ownership is clear because ownership is total: each team owns everything they build.</p><p>A platform team forms. Its first mandate is to centralize shared services: CI/CD, networking, and account management. It takes those over. But app teams keep their existing infrastructure provisioning rights. Nobody wants to slow them down.</p><p>The platform team grows. It takes on reliability objectives. It takes on a compliance scope. It takes on cost accountability. Each expansion is reasonable in isolation.</p><p>What nobody updates is the boundary. App teams still have full autonomy over their own infrastructure. The platform now has accountability for the consequences of that autonomy.</p><p>By the time this becomes visible, it is embedded. The platform team is measured against outcomes they cannot fully control. Their performance review includes metrics that depend on decisions they have no input into.</p><div><hr></div><h2>The Operating Principle That Resolves This</h2><p>Accountability must match authority. This is not an organizational theory. It is an operating principle that can be implemented in weeks.</p><p>The practical form: <strong>platform teams should have approval rights over any infrastructure decision that touches their accountability scope.</strong></p><p>That means:</p><ul><li><p>If the platform owns the cloud cost review, the platform approves compute and storage provisioning above the defined thresholds</p></li><li><p>If the platform owns the compliance posture, the platform reviews and approves infrastructure changes that affect controls in scope</p></li><li><p>If the platform owns the reliability objective, the platform sets the change management policy that governs deployments to production</p></li></ul><p>This does not require the platform to do all the work. App teams still build. They still deploy. The platform provides the standards, the guardrails, and, in defined cases, the approval gate.</p><p>The goal is not control. The goal is alignment between who owns the risk and who influences the decisions that create it.</p><div><hr></div><h2>What Improves When You Get This Right</h2><p>When accountability matches authority, three things stabilize quickly.</p><p>Incidents become attributable. When platform standards govern the infrastructure, post-incident reviews identify gaps in the standard, not just in the team that missed it. The root cause analysis becomes systemic. The fix improves the platform, not just the individual response.</p><p>Cost reviews become credible. When the platform controls tagging policy and provisioning standards, attribution improves. You can present AWS spend by team, by product, or by tenant. The CFO gets a useful number. The CTO can make decisions from it.</p><p>Audit prep becomes predictable. When the platform owns the change classification model and the approval gates, control coverage is tracked continuously. Evidence collection becomes operational rather than reactive.</p><p>The platform team stops being the team that explains what went wrong. It becomes the team that designed the system that prevented it.</p><div><hr></div><p><em>On Wednesday, paid subscribers get the full operating model for implementing this: the approval decision tree, the change classification model, and the sensitive-change checklist I would use to define which infrastructure decisions need platform review, which can be self-service, and which must be blocked.</em></p><h2><strong>That&#8217;s it for today!</strong></h2><p>Did you enjoy this newsletter issue?</p><p>Share with your friends, colleagues, and your favorite social media platform.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Cloud Playbook</span></a></p><p><strong>Until next week &#8212; Amrut</strong></p><div><hr></div><h2><strong>Get in touch</strong></h2><p>You can find me on <a href="https://www.linkedin.com/in/patilamrut/">LinkedIn</a> or <a href="https://twitter.com/realamrutpatil">X</a>.</p><p>If you would like to request a topic to read, please feel free to contact me directly via LinkedIn or X.</p>]]></content:encoded></item><item><title><![CDATA[TCP #120: The Infrastructure Ownership Matrix For Platform And App Teams]]></title><description><![CDATA[A practical way to decide who can provision which AWS resources, under what conditions, and with whose approval.]]></description><link>https://www.thecloudplaybook.com/p/infrastructure-ownership-matrix-platform-app-teams-aws</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/infrastructure-ownership-matrix-platform-app-teams-aws</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Wed, 06 May 2026 11:00:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!QX_P!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2488e5be-47fd-437b-bec0-b260eafc96cb_1086x1448.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Sunday&#8217;s issue named the problem: infrastructure provisioned without an ownership model creates reliability gaps, cost exposure, and audit risk.</p><p>This issue gives you the system.</p><p><strong>The Infrastructure Ownership Matrix</strong> defines who can provision what, under what conditions, and who approves what changes. It replaces the informal agreements that break down as soon as a team grows or a new engineer joins.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QX_P!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2488e5be-47fd-437b-bec0-b260eafc96cb_1086x1448.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QX_P!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2488e5be-47fd-437b-bec0-b260eafc96cb_1086x1448.png 424w, https://substackcdn.com/image/fetch/$s_!QX_P!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2488e5be-47fd-437b-bec0-b260eafc96cb_1086x1448.png 848w, https://substackcdn.com/image/fetch/$s_!QX_P!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2488e5be-47fd-437b-bec0-b260eafc96cb_1086x1448.png 1272w, https://substackcdn.com/image/fetch/$s_!QX_P!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2488e5be-47fd-437b-bec0-b260eafc96cb_1086x1448.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QX_P!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2488e5be-47fd-437b-bec0-b260eafc96cb_1086x1448.png" width="1086" height="1448" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2488e5be-47fd-437b-bec0-b260eafc96cb_1086x1448.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1448,&quot;width&quot;:1086,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1426294,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/196369637?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2488e5be-47fd-437b-bec0-b260eafc96cb_1086x1448.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QX_P!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2488e5be-47fd-437b-bec0-b260eafc96cb_1086x1448.png 424w, https://substackcdn.com/image/fetch/$s_!QX_P!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2488e5be-47fd-437b-bec0-b260eafc96cb_1086x1448.png 848w, https://substackcdn.com/image/fetch/$s_!QX_P!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2488e5be-47fd-437b-bec0-b260eafc96cb_1086x1448.png 1272w, https://substackcdn.com/image/fetch/$s_!QX_P!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2488e5be-47fd-437b-bec0-b260eafc96cb_1086x1448.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why This Problem Gets Expensive Before Anyone Notices</h2>
      <p>
          <a href="https://www.thecloudplaybook.com/p/infrastructure-ownership-matrix-platform-app-teams-aws">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[TCP #119: Platform Teams Do Not Scale by Saying Yes Faster]]></title><description><![CDATA[Most platform bottlenecks come from unclear intake, routing, approvals, and ownership, not a lack of headcount.]]></description><link>https://www.thecloudplaybook.com/p/platform-team-bottlenecks-intake-routing-approvals</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/platform-team-bottlenecks-intake-routing-approvals</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Mon, 04 May 2026 11:01:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!RGNt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736ef2b7-1141-480a-90bf-8980fb64472b_1492x1054.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Requests pile up. Developers escalate to their managers, who escalate to platform leadership.</p><p>The SLA misses compound. Engineers work hard and still fall behind.</p><p>Every VP who sees this situation reaches the same conclusion: the platform team needs more headcount.</p><p>That conclusion is almost always wrong.</p><p>Platform team bottlenecks do not come from teams that are too small. They come from work, arrive through unclear intake channels, are routed to ambiguous owners, and wait for approvals nobody documented.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RGNt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736ef2b7-1141-480a-90bf-8980fb64472b_1492x1054.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RGNt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736ef2b7-1141-480a-90bf-8980fb64472b_1492x1054.png 424w, https://substackcdn.com/image/fetch/$s_!RGNt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736ef2b7-1141-480a-90bf-8980fb64472b_1492x1054.png 848w, https://substackcdn.com/image/fetch/$s_!RGNt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736ef2b7-1141-480a-90bf-8980fb64472b_1492x1054.png 1272w, https://substackcdn.com/image/fetch/$s_!RGNt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736ef2b7-1141-480a-90bf-8980fb64472b_1492x1054.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RGNt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736ef2b7-1141-480a-90bf-8980fb64472b_1492x1054.png" width="1456" height="1029" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/736ef2b7-1141-480a-90bf-8980fb64472b_1492x1054.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1029,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1454431,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/196368219?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736ef2b7-1141-480a-90bf-8980fb64472b_1492x1054.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RGNt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736ef2b7-1141-480a-90bf-8980fb64472b_1492x1054.png 424w, https://substackcdn.com/image/fetch/$s_!RGNt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736ef2b7-1141-480a-90bf-8980fb64472b_1492x1054.png 848w, https://substackcdn.com/image/fetch/$s_!RGNt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736ef2b7-1141-480a-90bf-8980fb64472b_1492x1054.png 1272w, https://substackcdn.com/image/fetch/$s_!RGNt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736ef2b7-1141-480a-90bf-8980fb64472b_1492x1054.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Why Faster Ticket Response Does Not Fix the Platform Engineering Bottleneck</h2><p>The reflexive response to a growing platform team backlog is to optimize throughput.</p><p>Run intake meetings twice a week instead of once. Add a triage rotation. Write SLA targets. Bring in a TPM to route requests. Some leaders introduce a tiered priority system: P0 gets a 24-hour response, P1 gets a five-day response, and P2 gets a two-week response.</p><p>Each change makes the intake process marginally more efficient. None of them fixes what actually causes requests to stall.</p><p>They do not tell a developer where to submit a request when their Slack message from two weeks ago went unanswered. They do not clarify which approval an engineer needs to unblock a security exception. They do not identify who owns an ambiguous request when it lands in the queue with no routing context.</p><p>Adding a priority label to an unrouted request does not route it.</p><p>Faster throughput into an unclear structure is still unclear.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Cloud Playbook is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Four Structural Gaps That Make Platform Teams a Bottleneck</h2><p>Platform engineering scalability problems follow a consistent pattern. Four things are missing at once.</p><ol><li><p><strong>Intake clarity.</strong> There is no single, well-defined path to request platform work. Some teams submit tickets. Others use Slack. Some skip both and go straight to a platform engineer.</p><p>Because the platform team's intake process is informal, everything arrives marked urgent. The team cannot distinguish a genuine blocker from a request that can wait two weeks.</p></li><li><p><strong>Routing clarity.</strong> Once a request lands, no one is certain who will handle it. The team is large enough that ownership is ambiguous.</p><p>Requests get forwarded, sit in limbo, or wait for whoever happens to know the most about that area. There is no platform team request routing logic written down anywhere.</p></li><li><p><strong>Approval clarity.</strong> New infrastructure, security exceptions, and networking changes: each requires sign-off. But the approval chain is not documented.</p><p>Requests stall while engineers chase the right approver. Without a defined process, there is no predictable SLA for anything requiring sign-off, and every blocked request becomes a separate escalation path.</p></li><li><p><strong>Ownership clarity.</strong> When something breaks or a decision needs to be made, &#8220;Who owns this?&#8221; takes too long to answer.</p><p>If developer platform ownership is ambiguous during normal operations, it becomes a crisis under pressure. Every incident starts with a 20-minute conversation that should take 90 seconds.</p></li></ol><p>These four gaps appear to be a capacity problem from the outside. Inside, they feel like everyone is working hard, but nothing is moving.</p><p>Adding engineers to this structure does not fix it. It replicates it. Each new hire spends their first months navigating the same ambiguity the current team has learned to live with.</p><div><hr></div><h2>The Four Questions That Confirm a Structure Problem</h2><p>Before approving a headcount requisition, run this diagnostic.</p><ol><li><p>If a developer needs a new service account today, do they know exactly where to submit the request? Or does the answer depend on who they know?</p></li><li><p>When a request arrives, can your platform engineer identify the owner in under five minutes without asking three colleagues?</p></li><li><p>For a security exception request, can you name the approver and the expected response time right now, without looking it up?</p></li><li><p>If you ask five engineers on your platform team, &#8220;Who owns the API gateway?&#8221; do you get the same answer within five minutes?</p></li></ol><p>One &#8220;it depends&#8221; in those answers means you have a platform team structure problem, not a headcount problem. Hiring more engineers will not change those answers.</p><div><hr></div><h2>How to Make Platform Team Structure Explicit Before You Hire</h2><p>These structural fixes cost less than a single hire and last longer than any retrospective.</p><p><strong>Define one intake channel.</strong> One Slack channel. One ticket form. One entry point for all requests.</p><p>Not &#8220;it depends on the request type.&#8221; One place. This makes the queue visible and eliminates the parallel-path problem where the same work gets started twice by two people who each received a slightly different version of the request.</p><p><strong>Build a routing matrix.</strong> For each request category, define who handles it by role, not name.</p><p>New service account: Platform Infrastructure team, reviewed Mondays. Security exception: Security guild plus Platform lead, SLA 5 business days. The matrix need not be complex. It needs to exist.</p><p><strong>Document the approval chain.</strong> For every request type requiring sign-off, name the role and the expected turnaround. Post it in your intake channel.</p><p>Approvals do not need to be fast. They need to be predictable.</p><p><strong>Assign single owners.</strong> Every platform component, every shared service, every critical decision needs one named person, not a team. Ownership rotates on a schedule. The clarity does not.</p><p>The goal is not to eliminate judgment from the platform team. It is to remove the structural overhead that consumes judgment before real work begins. When intake, routing, approvals, and ownership are clear, engineers spend more time engineering.</p><div><hr></div><h2>Run this check this week:</h2><p>Pull the last five platform requests that missed your SLA.</p><p>For each one, trace its entry into the system, its routing, who needed to approve it, and at which step it stopped moving.</p><p>That step is your structural gap. Fix it before opening a headcount requisition.</p><p>Teams that define intake, routing, and ownership before their next hire recover 30 to 40 percent of effective capacity without adding a single engineer. That is the capacity that the structural ambiguity was absorbing.</p><p>Every time I have traced a chronic platform team backlog to its root cause, the issue was structural: a missing routing matrix, an undocumented approval chain, and no one who could answer &#8220;who owns the API gateway&#8221; in under thirty seconds. </p><p>The team was not too small. The structure was invisible.</p><p><em>On Wednesday, paid subscribers get the full Platform Intake Operating Model: a routing matrix template, approval tiers, rollout checklist, and metrics you can implement with your team.</em></p><div><hr></div><h2>Upgrade If You Need Implementation, Not Just Ideas</h2><p>If you&#8217;re using these emails to guide real decisions on your platform, you&#8217;ll get more leverage from the paid version of The Cloud Playbook.</p><p>The free newsletter gives you patterns and language.</p><p>The paid newsletter turns those patterns into implementation kits you can ship inside a quarter:</p><ul><li><p>Concrete rollout plans (90&#8209;day roadmaps for each pattern)</p></li><li><p>Templates and checklists (policies, runbooks, tagging schemes, review checklists)</p></li><li><p>Real examples from high&#8209;stakes AWS environments (what we actually shipped and why)</p></li></ul><p>If the paid side doesn&#8217;t save you more than the subscription in <strong>one</strong> incident, audit cycle, or bad migration you avoid, you should cancel and keep the playbooks.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe&quot;,&quot;text&quot;:&quot;Upgrade to the Paid Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thecloudplaybook.com/subscribe"><span>Upgrade to the Paid Cloud Playbook</span></a></p><div><hr></div><h2><strong>That&#8217;s it for today!</strong></h2><p>Did you enjoy this newsletter issue?</p><p>Share with your friends, colleagues, and your favorite social media platform.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Cloud Playbook</span></a></p><p><strong>Until next week &#8212; Amrut</strong></p><div><hr></div><h2><strong>Get in touch</strong></h2><p>You can find me on <a href="https://www.linkedin.com/in/patilamrut/">LinkedIn</a> or <a href="https://twitter.com/realamrutpatil">X</a>.</p><p>If you would like to request a topic to read, please feel free to contact me directly via LinkedIn or X.</p>]]></content:encoded></item><item><title><![CDATA[TCP #118: The Platform Intake Operating Model For Scaling Platform Teams]]></title><description><![CDATA[How to replace Slack chaos with a routing matrix, approval tiers, and rollout plan your platform team can actually live with]]></description><link>https://www.thecloudplaybook.com/p/platform-intake-operating-model-routing-matrix-approvals-slas</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/platform-intake-operating-model-routing-matrix-approvals-slas</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Thu, 30 Apr 2026 15:03:04 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!VXPY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee53e291-8bf3-4c38-9f0b-a6de7dc69e68_1491x1055.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="https://open.substack.com/pub/thecloudplaybook/p/platform-team-bottleneck?r=ainou&amp;utm_campaign=post&amp;utm_medium=web">Sunday&#8217;s newsletter issue</a> made the case. Saying yes faster does not scale a platform team. Structure does.</p><p>When intake is informal, platform teams drown in tickets, senior engineers become human routers, and every request feels urgent. Reliability suffers because the work that prevents incidents gets displaced by whoever shouts the loudest.</p><p>A platform team that cannot control intake cannot control its outcomes.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VXPY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee53e291-8bf3-4c38-9f0b-a6de7dc69e68_1491x1055.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VXPY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee53e291-8bf3-4c38-9f0b-a6de7dc69e68_1491x1055.png 424w, https://substackcdn.com/image/fetch/$s_!VXPY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee53e291-8bf3-4c38-9f0b-a6de7dc69e68_1491x1055.png 848w, https://substackcdn.com/image/fetch/$s_!VXPY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee53e291-8bf3-4c38-9f0b-a6de7dc69e68_1491x1055.png 1272w, https://substackcdn.com/image/fetch/$s_!VXPY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee53e291-8bf3-4c38-9f0b-a6de7dc69e68_1491x1055.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VXPY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee53e291-8bf3-4c38-9f0b-a6de7dc69e68_1491x1055.png" width="1456" height="1030" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee53e291-8bf3-4c38-9f0b-a6de7dc69e68_1491x1055.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1030,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1777148,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/195580288?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee53e291-8bf3-4c38-9f0b-a6de7dc69e68_1491x1055.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VXPY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee53e291-8bf3-4c38-9f0b-a6de7dc69e68_1491x1055.png 424w, https://substackcdn.com/image/fetch/$s_!VXPY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee53e291-8bf3-4c38-9f0b-a6de7dc69e68_1491x1055.png 848w, https://substackcdn.com/image/fetch/$s_!VXPY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee53e291-8bf3-4c38-9f0b-a6de7dc69e68_1491x1055.png 1272w, https://substackcdn.com/image/fetch/$s_!VXPY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee53e291-8bf3-4c38-9f0b-a6de7dc69e68_1491x1055.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2><strong>The Human Router Platform</strong></h2><p>Most teams start with good intent.</p>
      <p>
          <a href="https://www.thecloudplaybook.com/p/platform-intake-operating-model-routing-matrix-approvals-slas">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[TCP #117: Your platform team doesn’t have a capacity problem.]]></title><description><![CDATA[4 structure checks to recover 30&#8211;40% of their time without hiring.]]></description><link>https://www.thecloudplaybook.com/p/platform-team-bottleneck</link><guid isPermaLink="false">https://www.thecloudplaybook.com/p/platform-team-bottleneck</guid><dc:creator><![CDATA[Amrut Patil]]></dc:creator><pubDate>Sun, 26 Apr 2026 14:21:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xxQU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4d86b01-1368-4adf-8e0b-e116e611887d_1122x1402.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Requests pile up. Developers escalate to their managers, who escalate to platform leadership.</p><p>The SLA misses compound. Engineers work hard and still fall behind.</p><p>Every VP who sees this situation reaches the same conclusion: the platform team needs more headcount.</p><p>That conclusion is almost always wrong.</p><p>Platform team bottlenecks do not come from teams that are too small. They come from work that arrives through unclear intake channels, gets routed to ambiguous owners, and waits for undocumented approvals.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xxQU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4d86b01-1368-4adf-8e0b-e116e611887d_1122x1402.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xxQU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4d86b01-1368-4adf-8e0b-e116e611887d_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!xxQU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4d86b01-1368-4adf-8e0b-e116e611887d_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!xxQU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4d86b01-1368-4adf-8e0b-e116e611887d_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!xxQU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4d86b01-1368-4adf-8e0b-e116e611887d_1122x1402.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xxQU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4d86b01-1368-4adf-8e0b-e116e611887d_1122x1402.png" width="1122" height="1402" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f4d86b01-1368-4adf-8e0b-e116e611887d_1122x1402.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1402,&quot;width&quot;:1122,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1492691,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thecloudplaybook.com/i/193221322?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4d86b01-1368-4adf-8e0b-e116e611887d_1122x1402.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xxQU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4d86b01-1368-4adf-8e0b-e116e611887d_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!xxQU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4d86b01-1368-4adf-8e0b-e116e611887d_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!xxQU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4d86b01-1368-4adf-8e0b-e116e611887d_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!xxQU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4d86b01-1368-4adf-8e0b-e116e611887d_1122x1402.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Why Faster Ticket Response Does Not Fix the Platform Engineering Bottleneck</h2><p>The reflexive response to a growing platform team backlog is to optimize throughput.</p><p>Run intake meetings twice a week instead of once. Add a triage rotation. Write SLA targets. Bring in a TPM to route requests. Some leaders introduce a tiered priority system: P0 gets a 24-hour response, P1 gets a five-day response, and P2 gets a two-week response.</p><p>Each change makes the intake process marginally more efficient. None of them fixes what actually causes requests to stall.</p><p>They do not tell a developer where to submit a request when their Slack message from two weeks ago went unanswered. They do not clarify which approval an engineer needs to unblock a security exception. They do not identify who owns an ambiguous request when it lands in the queue with no routing context.</p><p>Adding a priority label to an unrouted request does not route it.</p><p>Faster throughput into an unclear structure is still unclear structure.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Cloud Playbook is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Four Structural Gaps That Make Platform Teams a Bottleneck</h2><p>Platform engineering scalability problems follow a consistent pattern: four structural elements are usually missing at once.</p><ol><li><p><strong>Intake clarity.</strong> There is no single, well-defined path to request platform work. Some teams submit tickets. Others use Slack. Some skip both and corner a platform engineer directly. </p><p>Because the platform team intake process is informal, everything arrives marked urgent. The team cannot distinguish a genuine blocker from a request that can wait two weeks.</p></li><li><p><strong>Routing clarity.</strong> Once a request lands, no one is certain who will handle it. The team is large enough that ownership is ambiguous.</p><p>Requests get forwarded, sit in limbo, or wait for whoever happens to know the most about that area. There is no platform team request routing logic written down anywhere.</p></li><li><p><strong>Approval clarity.</strong> New infrastructure, security exceptions, and networking changes: each requires sign-off. But the approval chain is not documented.</p><p>Requests stall while engineers chase the right approver. Without a defined process, there is no predictable SLA for anything requiring sign-off, and every blocked request becomes a separate escalation path.</p></li><li><p><strong>Ownership clarity.</strong> When something breaks or a decision needs to be made, &#8220;Who owns this?&#8221; takes too long to answer. If developer platform ownership is ambiguous during normal operations, it becomes a crisis under pressure. Every incident starts with a 20-minute conversation that should take 90 seconds.</p></li></ol><p>These four gaps appear to be a capacity problem from the outside. Inside, they feel like everyone is working hard, but nothing is moving.</p><p>Adding engineers to this structure does not fix it. It replicates it. Each new hire spends their first months navigating the same ambiguity the current team has learned to live with.</p><div><hr></div><h2>The Four Questions That Confirm a Structure Problem</h2><p>Before approving a headcount requisition, run this diagnostic.</p><ol><li><p>If a developer needs a new service account today, do they know exactly where to submit the request? Or does the answer depend on who they know?</p></li><li><p>When a request arrives, can your platform engineer identify the owner in under five minutes without asking three colleagues?</p></li><li><p>For a security exception request, can you name the approver and the expected response time right now, without looking it up?</p></li><li><p>If you ask five engineers on your platform team, &#8220;Who owns the API gateway?&#8221; do you get the same answer within five minutes?</p></li></ol><p>One &#8220;it depends&#8221; in those answers means you have a platform team structure problem, not a headcount problem. Hiring more engineers will not change those answers.</p><div><hr></div><h2>How to Make Platform Team Structure Explicit Before You Hire</h2><p>These structural fixes cost less than a single hire and last longer than any retrospective.</p><ul><li><p><strong>Define one intake channel.</strong> One Slack channel. One ticket form. One entry point for all requests.</p><p>Not &#8220;it depends on the request type.&#8221; One place. This makes the queue visible and eliminates the parallel-path problem where the same work gets started twice by two people who each received a slightly different version of the request.</p></li><li><p><strong>Build a routing matrix.</strong> For each request category, define who handles it by role, not name.</p><p>New service account: Platform Infrastructure team, reviewed Mondays. Security exception: Security guild plus Platform lead, SLA 5 business days. The matrix need not be complex. It needs to exist.</p></li><li><p><strong>Document the approval chain.</strong> For every request type requiring sign-off, name the role and the expected turnaround. Post it in your intake channel.</p><p>Approvals do not need to be fast. They need to be predictable.</p></li><li><p><strong>Assign single owners.</strong> Every platform component, every shared service, every critical decision needs one named person, not a team. Ownership rotates on a schedule. The clarity does not.</p></li></ul><p>The goal is not to eliminate judgment from the platform team. It is to remove the structural overhead that consumes judgment before real work begins. When intake, routing, approvals, and ownership are clear, engineers spend more time engineering.</p><div><hr></div><h3>Run this check this week</h3><p>Pull the last five platform requests that missed your SLA.</p><p>For each one, trace its entry into the system, its routing, who needed to approve it, and at which step it stopped moving.</p><p>That step is your structural gap. Fix it before opening a headcount requisition.</p><p>Teams that define intake, routing, and ownership before their next hire recover 30 to 40 percent of effective capacity without adding a single engineer. That is the capacity that the structural ambiguity was absorbing.</p><p>Every time I have traced a chronic platform team backlog to its root cause, the issue was structural: a missing routing matrix, an undocumented approval chain, and no one who could answer &#8220;who owns the API gateway&#8221; in under thirty seconds. </p><p>The team was not too small. The structure was invisible.</p><div><hr></div><h3>Upgrade If You Need Implementation, Not Just Ideas</h3><p>If you&#8217;re using these emails to guide real decisions on your platform, you&#8217;ll get more leverage from the paid version of The Cloud Playbook.</p><p>The free newsletter gives you patterns and language.</p><p>The paid newsletter turns those patterns into implementation kits you can ship inside a quarter:</p><ul><li><p>Concrete rollout plans (90&#8209;day roadmaps for each pattern)</p></li><li><p>Templates and checklists (policies, runbooks, tagging schemes, review checklists)</p></li><li><p>Real examples from high&#8209;stakes AWS environments (what we actually shipped and why)</p></li></ul><p>If the paid side doesn&#8217;t save you more than the subscription in <strong>one</strong> incident, audit cycle, or bad migration you avoid, you should cancel and keep the playbooks.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/subscribe&quot;,&quot;text&quot;:&quot;Upgrade to the Paid Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/subscribe"><span>Upgrade to the Paid Cloud Playbook</span></a></p><div><hr></div><h2><strong>That&#8217;s it for today!</strong></h2><p>Did you enjoy this newsletter issue?</p><p>Share with your friends, colleagues, and your favorite social media platform.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Cloud Playbook&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.thecloudplaybook.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Cloud Playbook</span></a></p><p><strong>Until next week &#8212; Amrut</strong></p><div><hr></div><h2><strong>Get in touch</strong></h2><p>You can find me on <a href="https://www.linkedin.com/in/patilamrut/">LinkedIn</a> or <a href="https://twitter.com/realamrutpatil">X</a>.</p><p>If you would like to request a topic to read, please feel free to contact me directly via LinkedIn or X.</p>]]></content:encoded></item></channel></rss>