<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dave Kurian</title>
    <description>The latest articles on DEV Community by Dave Kurian (@davekurian).</description>
    <link>https://dev.to/davekurian</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3962819%2Ff0a481b6-456b-476e-bd6b-0aeb82ce4c1c.jpg</url>
      <title>DEV Community: Dave Kurian</title>
      <link>https://dev.to/davekurian</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/davekurian"/>
    <language>en</language>
    <item>
      <title>Buy the Arcade kit when the product is a games store</title>
      <dc:creator>Dave Kurian</dc:creator>
      <pubDate>Sun, 27 Sep 2026 13:41:20 +0000</pubDate>
      <link>https://dev.to/davekurian/buy-the-arcade-kit-when-the-product-is-a-games-store-38ba</link>
      <guid>https://dev.to/davekurian/buy-the-arcade-kit-when-the-product-is-a-games-store-38ba</guid>
      <description>&lt;p&gt;If the product you are shipping this quarter is a games store, buy the Arcade kit. The product page is a React Native and Expo template for that job: a five-tab store, hero carousels, an HTML5 game catalog, and a player profile, at $99. Browse, play, a library, and a player profile are the product. This is the SKU that already names those objects.&lt;/p&gt;

&lt;p&gt;The other kits you can buy alone today at $99 are the &lt;a href="https://otf-kit.dev/templates/saas-dashboard" rel="noopener noreferrer"&gt;SaaS Dashboard&lt;/a&gt; and the &lt;a href="https://otf-kit.dev/templates/fitness-kit" rel="noopener noreferrer"&gt;Fitness kit&lt;/a&gt;. Pricing lists those three as the live kits. Booking remains a preview on the bundle path until it is live. Marketplace is coming soon and is not sold on its own. Waiting on either one leaves the catalog, the library, and the profile unbuilt. If you already need two or more live kits in the same purchase, use &lt;a href="https://otf-kit.dev/pricing" rel="noopener noreferrer"&gt;pricing&lt;/a&gt; and the &lt;a href="https://otf-kit.dev/blog/everything-bundle-vs-single-kit" rel="noopener noreferrer"&gt;pack-versus-one-kit post&lt;/a&gt;. If the only product is the games store, checkout is the Arcade kit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Name the product before you open checkout
&lt;/h2&gt;

&lt;p&gt;A games store, for this decision, is a specific object. People open a home, scan a featured row, browse by genre, open a game, play it, keep a library, and return to a profile that remembers them. Search, filters, and wishlists sit on that path. Sign-in exists so the library belongs to someone.&lt;/p&gt;

&lt;p&gt;Catalog, play, library, and profile means this SKU. A dashboard product means the SaaS Dashboard. A fitness product means the Fitness kit. All three live kits are $99. The product object is the difference. Booking and Marketplace appear beside them as other cards. Pricing does not sell those two as solo live kits, so they stay off this checkout.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you already have on day one
&lt;/h2&gt;

&lt;p&gt;The Arcade product page is the inventory. The hero describes one template: React Native and Expo, a five-tab games store, hero carousels, a real HTML5 game catalog, and a player profile.&lt;/p&gt;

&lt;p&gt;Discover and play is a dark-first home with a featured hero carousel, genre browsing, full game-detail pages, and instant search. The labels in that block are Home, Store, Browse by Genre, Game Detail, and Search.&lt;/p&gt;

&lt;p&gt;The library block is instant-play picks, the player's game library, and a profile with level, coins, and an Arcade Pass state. Its labels are Instant Play, My Library, and Player Profile.&lt;/p&gt;

&lt;p&gt;Sign-in is a first-run onboarding flow plus email, Google, and guest, listed as wired. Its labels are Onboarding, Sign In, and Sign Up.&lt;/p&gt;

&lt;p&gt;The same page checks off hero carousels, a game store, a real game catalog, game detail pages, category browsing, search and filters, a player library, wishlists, and a player profile.&lt;/p&gt;

&lt;p&gt;The storefront on the page is seeded: genre chips for Action, Racing, Puzzle, Arcade, and Sports, a top-this-week list of titles, an Arcade Pass line, and a free-games count. That chrome is the UI. It is not your license count, and it is not a player-count claim.&lt;/p&gt;

&lt;p&gt;The stack block names Expo SDK 54, with one codebase for iOS, Android, and a web export; React Native 0.81; TypeScript 5.9; a Hono API; and Postgres with Drizzle, including schema, migrations, and seed data. The page says iOS and Android are tested. An Expo Go preview on SDK 54 opens the live kit with no build step.&lt;/p&gt;

&lt;p&gt;Every kit ships with CLAUDE.md, .cursorrules, and 26 tested prompts, aimed at Cursor, Claude, Lovable, and Bolt. That hands the repo to an assistant. You still read the catalog and the profile.&lt;/p&gt;

&lt;p&gt;The purchase line is one-time, source on GitHub, demo stays free, kit price $99.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0hstfh5k9lsz2quwk328.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0hstfh5k9lsz2quwk328.png" alt="Dex in purple gloves and Byte with bare hands compare a bright games-store tile to dimmer dashboard, fitness, and wait tiles while Luna with bare hands checks a short list; Nova floats above the bright tile." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Run this checklist before you pay
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;This quarter&lt;/th&gt;
&lt;th&gt;Checkout&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Games store only: browse, play, library, player profile&lt;/td&gt;
&lt;td&gt;Arcade kit, $99&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dashboard product only&lt;/td&gt;
&lt;td&gt;SaaS Dashboard template&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fitness product only&lt;/td&gt;
&lt;td&gt;Fitness kit template&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Two or more live kits, or a live kit plus a run of landings&lt;/td&gt;
&lt;td&gt;Pricing bundle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Booking or Marketplace as the games-store base&lt;/td&gt;
&lt;td&gt;Neither is a solo live kit, and neither is this catalog&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Stay on the Arcade row when the rest is also true.&lt;/p&gt;

&lt;p&gt;You would otherwise build the catalog, the library, and the profile before launch. You need iOS, Android, and a web export from the Expo SDK 54 codebase the page names. A marketing site on its own belongs with the landing templates. The &lt;a href="https://otf-kit.dev/blog/free-sdk-vs-paid-kit-when-to-buy" rel="noopener noreferrer"&gt;free SDK versus paid kit&lt;/a&gt; post is that fork. Day-one sign-in matches email, Google, or guest. You want the Hono API and the Postgres plus Drizzle setup the page names, including schema, migrations, and seed data. A server or database you have already committed to elsewhere is a port, so price that port before you pay.&lt;/p&gt;

&lt;p&gt;Pricing describes each live kit as login, payments, and a database already wired. The Arcade page names email, Google, and guest, plus Postgres. It does not print a payment-provider name in the stack list, so read that surface in the repo you receive.&lt;/p&gt;

&lt;p&gt;A second live kit already funded beside the games store moves the decision to &lt;a href="https://otf-kit.dev/pricing" rel="noopener noreferrer"&gt;pricing&lt;/a&gt; and the &lt;a href="https://otf-kit.dev/blog/everything-bundle-vs-single-kit" rel="noopener noreferrer"&gt;pack post&lt;/a&gt;. This kit is the games-store purchase. The pack is the purchase you make when a second live kit, or a stack of landings, is already on the plan.&lt;/p&gt;

&lt;h2&gt;
  
  
  The other live kits are different products
&lt;/h2&gt;

&lt;p&gt;Pricing's solo buy row is SaaS Dashboard, Fitness, and Arcade, each at $99.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://otf-kit.dev/templates/saas-dashboard" rel="noopener noreferrer"&gt;SaaS Dashboard template&lt;/a&gt; is the dashboard kit. Use it when that dashboard is the product. The &lt;a href="https://otf-kit.dev/templates/fitness-kit" rel="noopener noreferrer"&gt;Fitness kit template&lt;/a&gt; is the fitness product. Both can ship on a phone. A fitness product does not include the Arcade page's genre browse, playable HTML5 catalog, or profile with level, coins, and an Arcade Pass.&lt;/p&gt;

&lt;p&gt;The bundle post is where pack versus one kit gets decided. This post stays on the single games-store SKU.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4qfpfgkicit4gailze59.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4qfpfgkicit4gailze59.png" alt="Dex in purple gloves hands Byte with bare hands a single glowing Arcade kit phone-case box while Luna in a white research jacket closes a notebook on a wait pile; Nova floats with a soft cyan smile." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  When you do not need this kit
&lt;/h2&gt;

&lt;p&gt;Skip this kit when the product is a SaaS dashboard or a fitness app, and decide on those template pages.&lt;/p&gt;

&lt;p&gt;Skip it when you need components and themes, and the critical path has no catalog, library, sign-in, or database. Stay on the free SDK.&lt;/p&gt;

&lt;p&gt;Skip the single-SKU button when two or more live kits, or a live kit plus a run of landings, are already planned. The prices that govern the pack are printed on the pricing page.&lt;/p&gt;

&lt;p&gt;Skip the wait for a solo Booking or Marketplace buy button. Pricing marks Booking as a preview on the bundle path until it is live, and Marketplace as coming soon and not sold individually. A games catalog is not what those rows describe.&lt;/p&gt;

&lt;p&gt;Skip the purchase when you already run a catalog, accounts, and a profile service and only wanted a look at the UI. Open the demo. The kit is source for a store you intend to own.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you still do after the repo is yours
&lt;/h2&gt;

&lt;p&gt;The storefront is seed data rendered through the Drizzle schema: genres, a ranked list, search, and an Arcade Pass state. Replace that seed with the catalog you have rights to ship. Ratings, play counts, and free-game figures in the demo chrome stay demo figures until they are your numbers.&lt;/p&gt;

&lt;p&gt;Level, coins, and Arcade Pass are profile fields you can keep or rename. Coin sources, pass price, and what the demo's "no ads" and "cloud save" lines mean for your titles are still your call. The page shows those words on the pass line. Your economy is a later edit.&lt;/p&gt;

&lt;p&gt;Guest sign-in lets someone play before creating an account. You still decide how that session becomes an account and what happens to the library. New fields, wishlist rules, and the entitlement behind Arcade Pass land on the Hono API and the Drizzle schema. CLAUDE.md, .cursorrules, and the 26 prompts describe the repo to an assistant. Game licensing and the store listing stay with you.&lt;/p&gt;

&lt;p&gt;The versions printed on the page today are Expo SDK 54, React Native 0.81, and TypeScript 5.9. The Expo Go preview is pinned to SDK 54. Schedule any later upgrade on purpose.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check the demo, then use this checkout
&lt;/h2&gt;

&lt;p&gt;Open the live demo once at &lt;a href="https://arcade-preview.otf-kit.dev" rel="noopener noreferrer"&gt;https://arcade-preview.otf-kit.dev&lt;/a&gt; and walk home, the carousel, genre browse, a game detail, search, the library, and the profile. When that path is the product, buy it here: &lt;a href="https://otf-kit.dev/templates/arcade-kit" rel="noopener noreferrer"&gt;Arcade kit, $99&lt;/a&gt;. The page states a one-time purchase and source on GitHub.&lt;/p&gt;

&lt;p&gt;When the walk-through is a different product, use the checklist. The other live $99 pages are the SaaS Dashboard and the Fitness kit. When you need more than one live kit, the bundle decision stays on the pricing page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Arcade kit product page: &lt;a href="https://otf-kit.dev/templates/arcade-kit" rel="noopener noreferrer"&gt;https://otf-kit.dev/templates/arcade-kit&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Arcade demo: &lt;a href="https://arcade-preview.otf-kit.dev" rel="noopener noreferrer"&gt;https://arcade-preview.otf-kit.dev&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Pricing: &lt;a href="https://otf-kit.dev/pricing" rel="noopener noreferrer"&gt;https://otf-kit.dev/pricing&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;SaaS Dashboard template: &lt;a href="https://otf-kit.dev/templates/saas-dashboard" rel="noopener noreferrer"&gt;https://otf-kit.dev/templates/saas-dashboard&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Fitness kit template: &lt;a href="https://otf-kit.dev/templates/fitness-kit" rel="noopener noreferrer"&gt;https://otf-kit.dev/templates/fitness-kit&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Everything Bundle versus one kit: &lt;a href="https://otf-kit.dev/blog/everything-bundle-vs-single-kit" rel="noopener noreferrer"&gt;https://otf-kit.dev/blog/everything-bundle-vs-single-kit&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Free SDK versus a paid kit: &lt;a href="https://otf-kit.dev/blog/free-sdk-vs-paid-kit-when-to-buy" rel="noopener noreferrer"&gt;https://otf-kit.dev/blog/free-sdk-vs-paid-kit-when-to-buy&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>arcade</category>
      <category>expo</category>
      <category>reactnative</category>
    </item>
    <item>
      <title>Re-measure one fixed task before you trust a harness savings claim</title>
      <dc:creator>Dave Kurian</dc:creator>
      <pubDate>Sun, 27 Sep 2026 09:41:58 +0000</pubDate>
      <link>https://dev.to/davekurian/re-measure-one-fixed-task-before-you-trust-a-harness-savings-claim-dm9</link>
      <guid>https://dev.to/davekurian/re-measure-one-fixed-task-before-you-trust-a-harness-savings-claim-dm9</guid>
      <description>&lt;p&gt;After a coding-agent harness change ships, a SaaS team should re-measure one fixed task before it books a vendor percentage as its own bill. Keep a short standing file in the repo. Load each skill only when that task needs it. Run the same prompt on the same commit twice, once before the file change and once after, and keep a savings line only when the task still finishes.&lt;/p&gt;

&lt;p&gt;On 23 September 2026, Jediah Katz, Connor O'Keefe, and Calvin Yee published Cursor's account of harness work that reduced token costs for Cursor's users by 7% without reducing agent quality. That 7% is their result on their traffic. Your invoice moves when the same task, on your repo, uses fewer tokens and still finishes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Cursor changed in the harness
&lt;/h2&gt;

&lt;p&gt;Agents now run longer and carry context from step to step. Each turn resends tools, system instructions, setup, and the conversation, with a stable opening and a growing end. Cursor's production view splits spend into output, uncached input, and cached input. System and tool definitions in that view include compaction summaries. User text includes skills attached by hand. Skills and plugins include skill descriptions, MCP tool descriptions, and rules that sit in static context.&lt;/p&gt;

&lt;p&gt;Cursor controls the system prompt on every turn. As models improved, long lists of "do not," "you must," and "important" gave way to a shorter description of how a tool behaves. Cursor trimmed roughly 66% of that system prompt, across model families. They still add and remove lines when a new model needs guidance. They judge those edits with A/B tests on a large user base, because evals often represent hard problems and miss the true mix of requests.&lt;/p&gt;

&lt;p&gt;Tool definitions grew with background shell monitoring, cloud subagents, and more reliable web access. Most of those tools matter, and each is needed in fewer than 20% of conversations. Cursor had moved MCP tools into dynamic context earlier, which reduced total tokens by 46.9% on sessions that called an MCP tool. They applied the same pattern to built-in tools and A/B tested which ones stay visible from the first turn, watching tokens, cost, latency, tool-call errors, and agent usage. Reading, searching, editing, and the shell stayed in static context, along with &lt;code&gt;ask_question&lt;/code&gt;, because some models invented calls to it, and flow tools such as &lt;code&gt;create_plan&lt;/code&gt; in Plan Mode. The rest load when the agent needs them. That offload cut static-context description tokens by 60%.&lt;/p&gt;

&lt;p&gt;They then placed cache breakpoints after the stable layers and before the growing conversation. Before GPT-5.6 the cache boundary followed the latest request. Since GPT-5.6 the OpenAI API accepts explicit breakpoints beside its implicit cache. Cursor kept rarely changing tools and system instructions in front of that boundary, and moved skills, subagents, and environment info into a "phantom user message." Those cache changes reduced the rate of cold cache misses by 20%.&lt;/p&gt;

&lt;p&gt;The Read tool now numbers every tenth line. One line number uses around three to five tokens, and numbering every tenth line reduced cache-read tokens by 1.6% with no reduction in quality. Cursor also dropped instructions that strongly pushed subagents for codebase exploration, and a subagent switches model only when the user or the harness directs it. Those edits stay in their harness.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the 7% stays on their traffic
&lt;/h2&gt;

&lt;p&gt;The 7% is the combined result on Cursor's users, with no reduction in agent quality. The inner figures use other denominators. The 66% is Cursor's system prompt. The 60% is static tool-description tokens after the built-in-tool offload. The 20% is the rate of cold cache misses. The 46.9% is total tokens on sessions that called an MCP tool. The 1.6% is cache-read tokens from the Read tool's line numbers. Keep each figure on its denominator. Your figure is the gap between two finished runs of one task you already ship.&lt;/p&gt;

&lt;p&gt;Evals on Cursor's side often represent hard problems and miss everyday requests. A one-off hard task has the same limit, so pick a task from a normal week. Hold the model fixed. Model choice is a separate decision, in &lt;a href="https://otf-kit.dev/blog/cursor-router-ai-costs" rel="noopener noreferrer"&gt;how to route a coding agent across models&lt;/a&gt;. Hold the effort setting fixed too, as in &lt;a href="https://otf-kit.dev/blog/reduce-ai-token-costs" rel="noopener noreferrer"&gt;the effort knob on a coding agent&lt;/a&gt;. When the model, the effort setting, or the commit changes between runs, label the pair mixed and rerun with those three held still.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep a short standing file
&lt;/h2&gt;

&lt;p&gt;A long standing file rides along on tasks that never use most of it. Keep that file short. Put each skill in its own file and read it only for the matching task. Cursor put skills in the phantom user message so variable setup sat past the cache boundary. The block below is a team file shape. Cursor left the repo layout to the team.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# agent/standing.md
# Short on purpose. Change it in a reviewed diff.
# Keep skill bodies in their own files.

Run the app with the command in the README.
The fixed billing task is: add a usage row for the signed-in workspace and show it on the dashboard.
Tests for that task are the usage tests named in the README.
Edit the usage path and its test. Leave payment webhooks on their current path.
When the task is the usage row, read agent/skills/usage-row.md.
For any other task, leave that skill unread.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A reviewer should read this file in one sitting. &lt;a href="https://otf-kit.dev/blog/claude-md-under-200-lines" rel="noopener noreferrer"&gt;Keep project instructions that short&lt;/a&gt;. A skill body inside this file makes every later task pay for one task's steps.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcct3e5youwb2jj22hbyd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcct3e5youwb2jj22hbyd.png" alt="Byte edits a short standing file while Dex pins a thin stable layer card and Luna points" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Load the skill only for that task
&lt;/h2&gt;

&lt;p&gt;The standing file names the path. The agent reads it for the usage-row task and skips it otherwise. That follows the same idea as Cursor loading a built-in tool when the turn needs it, using files in your repo.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# agent/skills/usage-row.md
# Read this only when the task is the usage row.

Read the workspace id from the session the app already uses.
Write one usage row for that workspace.
Show the row on the dashboard page that already lists usage.
If the workspace id is missing, stop and leave the workspace list unchanged.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Edit the skill in place. The next run that reads the path sees the diff. Point the agent at that path like any other file in the repo. There is no extra service to call. Leave one short rule in the standing file when every task needs it. Move a second skill out when unrelated tasks would carry it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure the same task twice
&lt;/h2&gt;

&lt;p&gt;Use one prompt, one commit, and one model. Start each run in a fresh chat. Record only fields the product shows.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# agent/checks/usage-row-remeasure.txt
# Fill both columns from the product. Leave an unknown cell blank.

prompt: (paste the exact task prompt)
commit:
model:
effort setting, if the product has one:

                before                 after
standing file:  previous copy          short file + skill path
task finished:
input tokens:
cached input:
uncached input:
output tokens:
cost shown:
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faqe3ajw4zp7c2fkse6oq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faqe3ajw4zp7c2fkse6oq.png" alt="Dex and Luna complete a before-after checklist while Byte files it in the repo" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Completion comes first. A lower token count is a candidate saving only when the after run still finishes. Cursor paired their 7% with no reduction in agent quality. Your check is this task's diff and its test.&lt;/p&gt;

&lt;p&gt;Fill token cells from the product. When you only have a monthly invoice, write "not shown" and note the invoice period. A monthly total that mixes other people and other tasks stays a monthly total. Change one thing between the columns: shorten the standing file, or move one skill out of it. When a Cursor update lands in the same window, rerun after it settles, with the file as the only change.&lt;/p&gt;

&lt;h2&gt;
  
  
  When you don't need this
&lt;/h2&gt;

&lt;p&gt;Skip the files when this repo has no coding agent. The post is still worth reading. The re-measure is for a team about to book a savings line.&lt;/p&gt;

&lt;p&gt;Skip the pair while this repo still runs the build from before the harness change. Skip a second pair when someone already recorded two runs of this prompt and the standing file is unchanged. Run again when the next sprint would copy 7% onto a forecast.&lt;/p&gt;

&lt;p&gt;Skip the skill split when one short rule serves every task. Leave Cursor's tool list, cache breakpoints, and tenth-line numbering in their harness. The 60%, the 20%, and the 1.6% stay on the denominators in their post.&lt;/p&gt;

&lt;p&gt;Commit the standing file, the skill file, and the re-measure note together. Fill the before column from a run you trust. Run the after column in a fresh chat at that commit. Write a savings line only after both runs finish and the token cells come from the product.&lt;/p&gt;

&lt;p&gt;Ship the usage surface that task edits from the SaaS dashboard template: &lt;a href="https://otf-kit.dev/templates/saas-dashboard" rel="noopener noreferrer"&gt;https://otf-kit.dev/templates/saas-dashboard&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The template is the product shell for that dashboard. You add the standing file and the skill in the repo you already ship.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://cursor.com/blog/improved-token-efficiency" rel="noopener noreferrer"&gt;https://cursor.com/blog/improved-token-efficiency&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>agents</category>
      <category>cursor</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Pin a short skill allowlist instead of shopping the registry</title>
      <dc:creator>Dave Kurian</dc:creator>
      <pubDate>Sun, 27 Sep 2026 07:45:35 +0000</pubDate>
      <link>https://dev.to/davekurian/pin-a-short-skill-allowlist-instead-of-shopping-the-registry-511l</link>
      <guid>https://dev.to/davekurian/pin-a-short-skill-allowlist-instead-of-shopping-the-registry-511l</guid>
      <description>&lt;p&gt;A SaaS team should pin a short agent-skills allowlist in the repo instead of shopping the public registry every sprint. Amelia Charles, Andrew Qu, and Jonathan Hefner published the report behind that decision on the Vercel blog on 25 September 2026. The useful finding is concentration. Nearly half of all skills on the skills.sh registry were installed exactly once. 375 skills, 0.04% of the registry, account for 62% of installs. The most-installed skill is still under 1% of installs. There is no default skill to buy.&lt;/p&gt;

&lt;p&gt;Install figures in that report are aggregate registry counters. They do not represent unique people, and they do not necessarily represent independent choices. A high count can be one script repeating an install. A count of one can be a skill that fit a single setup. Neither count shows that an agent does the job better with the skill than without it. Pin the public skills that match jobs this repo already runs, and write the private rule only your company knows. The report names the shape of that rule: when a customer gets a refund, or what can ship without another review. Those are files you maintain, not rows in the registry.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the install curve actually says
&lt;/h2&gt;

&lt;p&gt;The top 1.2% of skills account for 94% of installs. Skills installed exactly once sit in what remains.&lt;/p&gt;

&lt;p&gt;That concentration is not one winner. Skills compete inside a job and stack across jobs. Once a team has an expense-report skill, it has little reason to install a second one for the same job. It can still add a different skill for spreadsheets, research, or a presentation. Each job has a small winner, and those winners together carry most installs. The allowlist is that set of jobs, written in the repo.&lt;/p&gt;

&lt;p&gt;A leaderboard is a poor source for the file. First place can be a job you do not run, and even that skill is under 1% of installs. Catalog totals count unique listings, and the report notes that figures may be revised as the data and the method improve. Treat the percentages as the shape on 25 September 2026. Treat the file as the decision that survives a revision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why an install count is a weak buy signal
&lt;/h2&gt;

&lt;p&gt;The report says the best public signal of quality today is install count. It is still a counter, not a test that the job got better. A skill justifies itself only when the agent does the job better with the skill than without it, and the report expects skills to gain tests for that. Until a test exists for the listing you want to add, the count cannot close the review.&lt;/p&gt;

&lt;p&gt;Category shares describe a classified sample, not every listing. No category draws more than a fifth of installs. Software engineering is the largest share at 18%, agent workflows follow at 15%, and business operations and writing sit near 11% each. None of those shares names a package.&lt;/p&gt;

&lt;p&gt;In that same sample, skills for work shared across industries are 66% of classified skills and 87.5% of installs. They draw 3.6 times as many installs per listing as industry-specific skills. That supports a short list of shared jobs. It does not let a public listing stand in for your refund rule.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pin the allowlist in the repo
&lt;/h2&gt;

&lt;p&gt;The pin is a file you review. The report does not define a command, an endpoint, or a schema for an allowlist, so this draft does not invent one. One job per block. The registry id stays blank until a reviewer has opened that listing. The date is the merge date. A second block for the same job needs a reason the first skill cannot do the work.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# agent/skills-allowlist.txt
# One public skill per job. Add a block only in a reviewed change.
# Install counts are aggregate counters, not unique people.

job: expense-report
registry-id:
pinned: 2026-09-27
note: one skill for this job

job: spreadsheet-cleanup
registry-id:
pinned: 2026-09-27
note: tidy an export before a person reads it
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those job names follow the report's examples. They are not packages it ranked. A blank &lt;code&gt;registry-id&lt;/code&gt; is not approved for install. Fill it in the same change that reviewed the listing. Do not paste an id from a leaderboard nobody opened.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzecqh2oy91kc0yoixoiv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzecqh2oy91kc0yoixoiv.png" alt="Byte edits the allowlist file while Luna points at two job blocks and Dex holds one approved card" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The pull request should name the weekly job, say why the current block is not enough if one exists, and say what a later review will re-check. If it cannot name a job in this repo, it does not merge. The next sprint reads the file. It does not browse the registry.&lt;/p&gt;

&lt;p&gt;Delete a block unused for a month. On sprint day, a filled id does not get a replacement search. A blank id gets one reviewed listing, or the job is marked private. A new job starts as a blank block. The install happens after the merge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Write the private rule next to the list
&lt;/h2&gt;

&lt;p&gt;The report says public skills will turn general know-how into the baseline. More of what sets a company apart will then sit in judgment only that company has. The examples it names are when a customer gets a refund, and what can ship without another review. Write them in ordinary language: the steps, what good looks like, and when to stop. The report describes a skill as a file of that kind. The registry cannot rank yours, because it does not know your window or which paths touch billing.&lt;/p&gt;

&lt;p&gt;The numbers below are a shape, not thresholds from the report, and not a default for every team. Replace each one with the rule a person on your team already follows.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# agent/skills/company-rules.md

When a refund is allowed

Read the charge on the current workspace only.
Allow the refund only when every line is true:
- The charge posted within 14 days.
- This workspace has no refund recorded in the current quarter.
- The caller is an owner of this workspace.
If any line is false, stop. Write the failed line on the ticket and leave the charge.

What can ship without another review

A change may ship with no extra reviewer only when every line is true:
- It does not change a price, a limit, or who can sign in.
- It touches no migration and no billing file.
- The diff is at most 40 lines.
If any line is false, open a review and stop. Do not merge.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq55nxtticunl4ynk51qv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq55nxtticunl4ynk51qv.png" alt="Dex and Luna review a private refund and ship-gate rule folder while Byte files it in the repo" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Change the rule by editing the file. A window that moves from 14 days to 7 days is a diff, and the agent sees it the next time it reads the path. Point the agent at that file the same way you point it at any other instruction file in the repo. There is no extra service to call.&lt;/p&gt;

&lt;p&gt;Keep three cases beside the file: a refund that meets every line, a refund that fails one line, and a diff that must wait. Re-run them when the file changes. That check stays private. It is not a registry count. Once a public skill is on the allowlist, &lt;a href="https://otf-kit.dev/blog/claude-code-plugin-eval-ci" rel="noopener noreferrer"&gt;measure whether that skill does the job&lt;/a&gt; before you add another beside it.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the catalog is already finite
&lt;/h2&gt;

&lt;p&gt;Some tools ship a small catalog for their own surface. Installing it once is a closed list. &lt;a href="https://otf-kit.dev/blog/expo-skills-for-ai-agents" rel="noopener noreferrer"&gt;Install that catalog&lt;/a&gt; and record those jobs as blocks whose note names the vendor catalog as the source, so the next sprint does not add a second public skill for a covered job.&lt;/p&gt;

&lt;p&gt;Do not paste that catalog into the private rule. It knows the tool. It does not know when your company refunds a charge. If a job still has a blank registry id, review one listing or keep the job private.&lt;/p&gt;

&lt;h2&gt;
  
  
  When you don't need this
&lt;/h2&gt;

&lt;p&gt;Skip the allowlist when no agent runs against this repo. The concentration in the report matters when someone is about to install a skill.&lt;/p&gt;

&lt;p&gt;Skip it when one person installs skills and already keeps that list in a note they open. The pin matters when the next person, or the next sprint, would shop again and treat a fresh count as a reason to switch.&lt;/p&gt;

&lt;p&gt;Skip the refund lines and the ship lines when a person still makes every call and the agent is not allowed to take them. Write the file when the agent is about to act.&lt;/p&gt;

&lt;p&gt;Skip a second public skill for a job that already has a block. Add a block for a new job instead. You also do not need the file merely to read the report.&lt;/p&gt;

&lt;p&gt;Put both files in one reviewed change. The list records which public skills are allowed. The rule records what only your company knows. Commit them in the SaaS repo you already ship. The SaaS dashboard template is here: &lt;a href="https://otf-kit.dev/templates/saas-dashboard" rel="noopener noreferrer"&gt;https://otf-kit.dev/templates/saas-dashboard&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://vercel.com/blog/state-of-agent-skills" rel="noopener noreferrer"&gt;https://vercel.com/blog/state-of-agent-skills&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.skills.sh/" rel="noopener noreferrer"&gt;https://www.skills.sh/&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>agents</category>
      <category>vercel</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Delete one tenant's files without scanning a shared Blob store</title>
      <dc:creator>Dave Kurian</dc:creator>
      <pubDate>Sat, 26 Sep 2026 15:45:30 +0000</pubDate>
      <link>https://dev.to/davekurian/delete-one-tenants-files-without-scanning-a-shared-blob-store-2dbc</link>
      <guid>https://dev.to/davekurian/delete-one-tenants-files-without-scanning-a-shared-blob-store-2dbc</guid>
      <description>&lt;p&gt;A multi-tenant app on Vercel can keep every customer's files in one Blob store under a pathname prefix, or give each customer a store. On 23 September 2026 Vercel removed the caps of 100 stores on Hobby, 500 on Pro, and 1,000 on Enterprise. One store per tenant is worth the create when export and delete must be a credential boundary. A mistaken pathname then cannot read or delete another tenant, because that credential does not open the other store. A prefix is still the right layout when your server builds every pathname from the authenticated tenant id and deletes from URLs saved on that tenant's rows.&lt;/p&gt;

&lt;p&gt;Creating a store is one Blob Advanced Operation, the same class as &lt;code&gt;put()&lt;/code&gt;, &lt;code&gt;copy()&lt;/code&gt;, and &lt;code&gt;list()&lt;/code&gt;. On Pro that class is $5.00 per million. On Hobby it counts toward the 2,000 advanced operations included each month. Deleting a store is free. Storage, operations, and data transfer for the same bytes cost the same in one store or in many. The create buys a unit you can hand to a job, wipe, or destroy. It does not buy a cheaper disk.&lt;/p&gt;

&lt;h2&gt;
  
  
  What still bills when stores are unlimited
&lt;/h2&gt;

&lt;p&gt;Storage size, simple operations, advanced operations, and data transfer still bill on use. Storage is the monthly average size. A simple operation is a cache-miss read by URL, or &lt;code&gt;head()&lt;/code&gt;. An advanced operation is &lt;code&gt;put()&lt;/code&gt;, &lt;code&gt;copy()&lt;/code&gt;, &lt;code&gt;list()&lt;/code&gt;, or a store create. Data transfer bills when a blob is downloaded. Every URL access is also an edge request, and a cache miss adds fast origin transfer. A blob larger than 512 MB is never cached, so every read of it is a miss.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;del()&lt;/code&gt; is free of charge. A batched &lt;code&gt;del([urls])&lt;/code&gt; still counts each blob toward the rate limit, so 100 pathnames are 100 operations against that limit, not one HTTP call. Hobby allows 1,200 simple and 900 advanced operations per minute. Pro allows 7,200 and 4,500. Dashboard browsing and uploads count as advanced operations too.&lt;/p&gt;

&lt;p&gt;Hobby includes 1 GB of storage, the first 10,000 simple operations, the first 2,000 advanced operations, and the first 10 GB of transfer each month. Past that, Blob access stops and the overage is not billed. Extra stores do not raise the 1 GB ceiling. About 2,000 creates in a month, with no &lt;code&gt;put()&lt;/code&gt; or &lt;code&gt;list()&lt;/code&gt; beside them, spend the Hobby advanced allowance on empty stores.&lt;/p&gt;

&lt;p&gt;Private and public stores share storage and operation prices. A private file is read by your server, then streamed to the caller, so both hops bill. A public URL is fetched by the client. Tenant documents belong in a private store. Access mode and region are fixed at create time. The CLI uses &lt;code&gt;iad1&lt;/code&gt; when you omit &lt;code&gt;--region&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  When a pathname prefix is enough
&lt;/h2&gt;

&lt;p&gt;One store, one credential. On Vercel the SDK pairs a short-lived OIDC token with one &lt;code&gt;BLOB_STORE_ID&lt;/code&gt; and refreshes that token. Outside Vercel, or for a signed browser upload, the credential is &lt;code&gt;BLOB_READ_WRITE_TOKEN&lt;/code&gt;. An explicit &lt;code&gt;token&lt;/code&gt; argument always beats OIDC.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwh3wpmooyl38r3e0tcwh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwh3wpmooyl38r3e0tcwh.png" alt="A shared store with pathname tabs is fragile when one credential opens every tenant" width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;del&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;list&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;put&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@vercel/blob&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;putTenantFile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Blob&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;put&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`tenants/&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;access&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;private&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;addRandomSuffix&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;deleteTenantPrefix&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;prefix&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`tenants/&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;undefined&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="k"&gt;do&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;prefix&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;blobs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;del&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;blobs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;blob&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;blob&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="nx"&gt;cursor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;hasMore&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cursor&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;undefined&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;addRandomSuffix&lt;/code&gt; defaults to false. Turn it on so two uploads of the same name do not collide, then delete the URL &lt;code&gt;put()&lt;/code&gt; returned. Each &lt;code&gt;list()&lt;/code&gt; page is one advanced operation. The default page size is 1,000. &lt;code&gt;del()&lt;/code&gt; does not throw if the URL is already gone, so a bad delete can look finished. Removed objects can stay on the CDN for up to a minute. Objects your rows forgot stay until a later list sees them.&lt;/p&gt;

&lt;p&gt;Keep the prefix when your server is the only writer, you never hand the credential to a tenant or another service, and tenants are small enough that the sweep is a few &lt;code&gt;list()&lt;/code&gt; calls. A leaked server credential that already exposes the rest of the account is not fixed by a second store. A public store fails even with a perfect prefix: anyone with the URL can read the file.&lt;/p&gt;

&lt;h2&gt;
  
  
  When a store per tenant is worth the create
&lt;/h2&gt;

&lt;p&gt;Create a user-created store when the unit you export, revoke, or destroy is the tenant. &lt;code&gt;POST https://api.vercel.com/storage/stores/blob&lt;/code&gt; takes &lt;code&gt;name&lt;/code&gt; (required, at most 70 characters), &lt;code&gt;access&lt;/code&gt; (&lt;code&gt;private&lt;/code&gt; or &lt;code&gt;public&lt;/code&gt;, default &lt;code&gt;public&lt;/code&gt;), and an optional &lt;code&gt;region&lt;/code&gt;. Set &lt;code&gt;access&lt;/code&gt; to &lt;code&gt;private&lt;/code&gt;. That POST is one advanced operation. &lt;code&gt;vercel blob create-store &amp;lt;name&amp;gt; --access private&lt;/code&gt; is the same create.&lt;/p&gt;

&lt;p&gt;The published 200 schema for that POST lists access, region, kind, size, and project metadata. It does not list an id, so do not invent &lt;code&gt;body.store.id&lt;/code&gt;. Match the name you sent with &lt;code&gt;vercel blob list-stores --all&lt;/code&gt;, save that id on the tenant row, and pass it as &lt;code&gt;storeId&lt;/code&gt;. Leave the project's single &lt;code&gt;BLOB_STORE_ID&lt;/code&gt; on the production store. Pointing it at a tenant store moves production traffic.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frl8ica8463ho0hfq0j5u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frl8ica8463ho0hfq0j5u.png" alt="One store credential is the unit you can hand out, empty, or delete" width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;put&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@vercel/blob&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;putInTenantStore&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;storeId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Blob&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;put&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;access&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;private&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;addRandomSuffix&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;storeId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;deleteTenantStore&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;storeId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;platformToken&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s2"&gt;`https://api.vercel.com/storage/stores/blob/&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nf"&gt;encodeURIComponent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;storeId&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;DELETE&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;Authorization&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Bearer &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;platformToken&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Blob store delete failed: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pass &lt;code&gt;storeId&lt;/code&gt; as &lt;code&gt;store_&amp;lt;id&amp;gt;&lt;/code&gt; or &lt;code&gt;&amp;lt;id&amp;gt;&lt;/code&gt;. Leave &lt;code&gt;VERCEL_OIDC_TOKEN&lt;/code&gt; in the environment. Copying it into &lt;code&gt;oidcToken&lt;/code&gt; skips refresh, and later calls fail with 403. A read-write &lt;code&gt;token&lt;/code&gt; is the override for a job outside Vercel, and it opens only that store.&lt;/p&gt;

&lt;p&gt;Read the tenant row before you POST. A retry opens a second store and bills another advanced operation. &lt;a href="https://otf-kit.dev/blog/idempotency-keys-owned-api" rel="noopener noreferrer"&gt;An owned API stores the first success and replays it&lt;/a&gt;. The first create is the one you keep.&lt;/p&gt;

&lt;p&gt;The platform token that may create and delete stores can destroy every tenant store. Keep it off the customer request path. &lt;a href="https://otf-kit.dev/blog/secrets-rotation-owned-backend" rel="noopener noreferrer"&gt;Rotate that secret with dual-key overlap&lt;/a&gt;. A project-default store is a different object: private, lazy, OIDC, and the mutation APIs do not manage it. Delete only a user-created store.&lt;/p&gt;

&lt;p&gt;Export is still &lt;code&gt;list()&lt;/code&gt; with no prefix, then &lt;code&gt;get()&lt;/code&gt; or the download URL. There is no export-store call. This &lt;code&gt;storeId&lt;/code&gt; cannot return the next tenant. Listing a large tenant still costs one advanced operation per thousand objects. Delete can skip that list.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;vercel blob empty-store &amp;lt;store-id&amp;gt;&lt;/code&gt; deletes every blob and keeps the store id. &lt;code&gt;vercel blob delete-store &amp;lt;store-id&amp;gt;&lt;/code&gt; removes the store. Both take &lt;code&gt;--yes&lt;/code&gt; in a script. Empty the store when the tenant stays. Delete it when the tenant is gone. Store deletion is free.&lt;/p&gt;

&lt;h2&gt;
  
  
  Split production, staging, and preview
&lt;/h2&gt;

&lt;p&gt;Production, staging, and shared preview are three stores. Three advanced operations, once, stop a preview write from landing on production objects. Make that split even when every production tenant shares one store.&lt;/p&gt;

&lt;p&gt;One preview store still mixes branches. Every deployment that shares &lt;code&gt;BLOB_STORE_ID&lt;/code&gt; can list the others. For a disposable branch, create a store, pass that &lt;code&gt;storeId&lt;/code&gt; only into that deployment, and &lt;code&gt;DELETE /storage/stores/blob/{id}&lt;/code&gt; when the branch closes. The delete is free. Bytes the preview wrote still bill until then. Do not point preview at the production store id to skip the create.&lt;/p&gt;

&lt;h2&gt;
  
  
  When you don't need this
&lt;/h2&gt;

&lt;p&gt;Skip a store per tenant for a single customer, or for public assets. Skip it when every blob URL already sits on the tenant row: offboarding is &lt;code&gt;del()&lt;/code&gt; of that list, and production still stays off the preview store. Skip it on Hobby if signups would create thousands of stores. About 2,000 creates consume the monthly advanced allowance. Stay on one private store and a server-built prefix until the plan can absorb them.&lt;/p&gt;

&lt;p&gt;Sweep a prefix as a paced job, not a tight loop. &lt;a href="https://otf-kit.dev/blog/queue-backpressure-owned-backend" rel="noopener noreferrer"&gt;Refuse that burst before it melts the rate limit&lt;/a&gt;. Each &lt;code&gt;list()&lt;/code&gt; page is an advanced operation, and each URL in a batched &lt;code&gt;del()&lt;/code&gt; counts toward the rate limit.&lt;/p&gt;

&lt;p&gt;OTF kits do not ship a per-tenant Blob layout yet. The owned upload path starts from the templates at &lt;a href="https://otf-kit.dev/templates" rel="noopener noreferrer"&gt;https://otf-kit.dev/templates&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://vercel.com/changelog/unlimited-vercel-blob-stores-on-every-plan" rel="noopener noreferrer"&gt;https://vercel.com/changelog/unlimited-vercel-blob-stores-on-every-plan&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vercel.com/docs/vercel-blob/usage-and-pricing" rel="noopener noreferrer"&gt;https://vercel.com/docs/vercel-blob/usage-and-pricing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vercel.com/docs/vercel-blob/using-blob-sdk" rel="noopener noreferrer"&gt;https://vercel.com/docs/vercel-blob/using-blob-sdk&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vercel.com/docs/rest-api/storage/create-a-blob-store" rel="noopener noreferrer"&gt;https://vercel.com/docs/rest-api/storage/create-a-blob-store&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vercel.com/docs/rest-api/storage/delete-a-blob-store" rel="noopener noreferrer"&gt;https://vercel.com/docs/rest-api/storage/delete-a-blob-store&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vercel.com/docs/cli/blob" rel="noopener noreferrer"&gt;https://vercel.com/docs/cli/blob&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vercel.com/kb/guide/vercel-blob" rel="noopener noreferrer"&gt;https://vercel.com/kb/guide/vercel-blob&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>vercel</category>
      <category>backend</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Write the outbox row in the same Postgres transaction as the business row</title>
      <dc:creator>Dave Kurian</dc:creator>
      <pubDate>Fri, 25 Sep 2026 09:34:21 +0000</pubDate>
      <link>https://dev.to/davekurian/write-the-outbox-row-in-the-same-postgres-transaction-as-the-business-row-3ljp</link>
      <guid>https://dev.to/davekurian/write-the-outbox-row-in-the-same-postgres-transaction-as-the-business-row-3ljp</guid>
      <description>&lt;p&gt;A welcome email, a webhook fan-out, and a search-index update are the same problem. The route has decided the business row. Another system still has to hear about it. Commit the user, then publish to the broker, and a broker that is down, or a process that dies, leaves a row with no event.&lt;/p&gt;

&lt;p&gt;That gap is a dual write. The database committed. The queue did not. Retrying the HTTP request can insert a second user, and it still may not send the first email. &lt;a href="https://otf-kit.dev/blog/idempotency-keys-owned-api" rel="noopener noreferrer"&gt;An owned API stores idempotency keys in Postgres and replays the first response&lt;/a&gt;. That row remembers the HTTP result. It does not publish an event you never wrote down.&lt;/p&gt;

&lt;p&gt;Write the outbox row in the same Postgres transaction as the business row. Commit both, or roll both back. A relay reads unpublished rows after commit and publishes them. The request does not talk to the broker. This is the &lt;a href="https://microservices.io/patterns/data/transactional-outbox.html" rel="noopener noreferrer"&gt;transactional outbox&lt;/a&gt;. The route is a Next.js handler. The database is Supabase Postgres.&lt;/p&gt;

&lt;h2&gt;
  
  
  The loss is the gap after commit
&lt;/h2&gt;

&lt;p&gt;Two writes that are not one transaction can succeed in either order, or only one of them.&lt;/p&gt;

&lt;p&gt;Publish first, then commit. The broker has the event and the transaction rolls back. Consumers act on a user who does not exist.&lt;/p&gt;

&lt;p&gt;Commit first, then publish. The user exists. The publish throws, times out, or never runs because the process was killed after commit. The client may already have a 201. Nothing in the database says the email is still owed.&lt;/p&gt;

&lt;p&gt;Holding the transaction open while you call the broker pins a session on the network. A broker error then rolls back a signup you were ready to keep. Do that only when you would rather fail the request.&lt;/p&gt;

&lt;p&gt;The outbox is the other choice. The business write must stick even if the broker is down, and the event must still leave later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Write the event in the same transaction
&lt;/h2&gt;

&lt;p&gt;The handler builds one unit of work: insert or update the business row, insert the outbox row, commit. supabase-js sends each statement as its own request. Those calls do not share a transaction. Put both writes in one Postgres function and call that function once, or use a client that holds a single transaction. If either insert fails, neither remains.&lt;/p&gt;

&lt;p&gt;Copy into &lt;code&gt;payload&lt;/code&gt; what the consumer must see as of this commit. A later update can change the user before the relay runs. A snapshot of the address belongs in the row. An event that means "read the user now" should say so in the type. Mixing the two indexes the wrong version.&lt;/p&gt;

&lt;p&gt;One event of a given type per aggregate gets a unique pair. A signup emits one &lt;code&gt;user.created&lt;/code&gt; for that user. A retried transaction hits the unique constraint and does not queue a second welcome. A stream of many events, such as every order change, uses its own id and does not use that pair.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fexg80qizd7ygldjm6gx7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fexg80qizd7ygldjm6gx7.png" alt="Write business and outbox rows in one transaction, then relay after commit" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The outbox table
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;published_at&lt;/code&gt; null means the relay still owes the broker this row. &lt;code&gt;claimed_at&lt;/code&gt; is a lease, not a receipt. &lt;code&gt;attempts&lt;/code&gt; counts how many times a relay has taken the row.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;create&lt;/span&gt; &lt;span class="k"&gt;table&lt;/span&gt; &lt;span class="n"&gt;outbox&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt; &lt;span class="k"&gt;primary&lt;/span&gt; &lt;span class="k"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;aggregate_type&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;aggregate_id&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;event_type&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="n"&gt;jsonb&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="n"&gt;timestamptz&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
  &lt;span class="n"&gt;published_at&lt;/span&gt; &lt;span class="n"&gt;timestamptz&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;claimed_at&lt;/span&gt; &lt;span class="n"&gt;timestamptz&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;attempts&lt;/span&gt; &lt;span class="nb"&gt;integer&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;unique&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;aggregate_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event_type&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;create&lt;/span&gt; &lt;span class="k"&gt;index&lt;/span&gt; &lt;span class="n"&gt;outbox_unpublished_idx&lt;/span&gt;
  &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="n"&gt;outbox&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;published_at&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;id&lt;/code&gt; is the event id the consumer will dedupe. Generate it in the route and store the same value in &lt;code&gt;payload&lt;/code&gt;. The partial index is the relay's queue. Published rows stay out of that index.&lt;/p&gt;

&lt;p&gt;The unique pair is the signup case. Drop it when the same event type can fire more than once, and keep &lt;code&gt;id&lt;/code&gt; as the primary key.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F38jnt6thkogweoglsi3w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F38jnt6thkogweoglsi3w.png" alt="Relay claims an unpublished outbox row with a lease while SKIP LOCKED skips a peer" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Poll with a claim, wake with notify
&lt;/h2&gt;

&lt;p&gt;The relay is a process that stays up. A Next.js route returns. It should not sit on &lt;a href="https://www.postgresql.org/docs/current/sql-listen.html" rel="noopener noreferrer"&gt;&lt;code&gt;LISTEN&lt;/code&gt;&lt;/a&gt;. &lt;code&gt;LISTEN&lt;/code&gt; needs a session that stays open. The durable record is the table. &lt;a href="https://www.postgresql.org/docs/current/sql-notify.html" rel="noopener noreferrer"&gt;&lt;code&gt;NOTIFY&lt;/code&gt;&lt;/a&gt; only wakes a relay that is already listening. In the default configuration the payload must be shorter than 8000 bytes, so it cannot carry the event. A notify sent while the relay is down is gone. The next poll still sees &lt;code&gt;published_at is null&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Claim a batch with &lt;a href="https://www.postgresql.org/docs/current/sql-select.html#SQL-FOR-UPDATE-SHARE" rel="noopener noreferrer"&gt;&lt;code&gt;FOR UPDATE SKIP LOCKED&lt;/code&gt;&lt;/a&gt; so two relays do not take the same row. The lease lets a crashed claim be taken again. Thirty seconds in the sketch is a lease you choose, not a measured optimum. Set &lt;code&gt;published_at&lt;/code&gt; after the broker accepts the message. A relay that dies after publish and before that update will publish again. That is at least once.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;due&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;
  &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;outbox&lt;/span&gt;
  &lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;published_at&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;
    &lt;span class="k"&gt;and&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;claimed_at&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt; &lt;span class="k"&gt;or&lt;/span&gt; &lt;span class="n"&gt;claimed_at&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;interval&lt;/span&gt; &lt;span class="s1"&gt;'30 seconds'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;order&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;
  &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="k"&gt;update&lt;/span&gt; &lt;span class="n"&gt;skip&lt;/span&gt; &lt;span class="n"&gt;locked&lt;/span&gt;
  &lt;span class="k"&gt;limit&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;update&lt;/span&gt; &lt;span class="n"&gt;outbox&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;
&lt;span class="k"&gt;set&lt;/span&gt; &lt;span class="n"&gt;claimed_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="n"&gt;attempts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;attempts&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;due&lt;/span&gt;
&lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;due&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;
&lt;span class="n"&gt;returning&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;event_type&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;attempts&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Polling on a timer is enough when a short delay is fine. &lt;code&gt;NOTIFY&lt;/code&gt; from the function that inserts the row, using the event id and not the body, shortens the wait when a listener is up. It does not replace the poll. A relay that only listens loses every event emitted while it was restarting.&lt;/p&gt;

&lt;h2&gt;
  
  
  At least once, then the consumer
&lt;/h2&gt;

&lt;p&gt;The broker and the consumer will see duplicates. Publish, crash, lease expires, publish again. The handler must treat &lt;code&gt;id&lt;/code&gt; as already done once it has run. Store that id the way an &lt;a href="https://otf-kit.dev/blog/idempotency-keys-owned-api" rel="noopener noreferrer"&gt;owned API stores an idempotency key and replays the first response&lt;/a&gt;. Record &lt;code&gt;id&lt;/code&gt; before the mailer call, or the second delivery sends a second email. How you hash a request body, and how long you keep the key, belong on that post.&lt;/p&gt;

&lt;p&gt;Do not set &lt;code&gt;published_at&lt;/code&gt; before the broker accepts the message. If the update that sets it fails, the duplicate is the consumer's problem. That is why the id was in the message.&lt;/p&gt;

&lt;p&gt;This table does not order events across aggregates. One relay can publish a single aggregate in &lt;code&gt;created_at&lt;/code&gt; order. Two relays, or parallel handlers, will not. If &lt;code&gt;user.updated&lt;/code&gt; must follow &lt;code&gt;user.created&lt;/code&gt;, the consumer serializes on &lt;code&gt;aggregate_id&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the relay keeps failing
&lt;/h2&gt;

&lt;p&gt;A payload the broker rejects, or a bug in the publisher, increments &lt;code&gt;attempts&lt;/code&gt; forever if every failure only drops the lease. Cap the attempts. Past the cap, stop selecting the row.&lt;/p&gt;

&lt;p&gt;Parking it, with the reason and enough of the payload to replay, is a dead-letter decision. &lt;a href="https://otf-kit.dev/blog/dead-letter-queue-background-jobs" rel="noopener noreferrer"&gt;Park failed jobs in a DLQ with the reason, the snapshot, and a replay&lt;/a&gt;. This post does not build that queue. The outbox's job ends when a person, or a replay tool, can see the event that did not leave.&lt;/p&gt;

&lt;p&gt;The relay is a loop and a publish. &lt;a href="https://otf-kit.dev/blog/ai-production-background-jobs" rel="noopener noreferrer"&gt;Background jobs for AI features&lt;/a&gt; covers queues and retries for that worker. Use the outbox when you must not lose the fact that the business row committed. Use an ordinary job when the caller may fail the request instead of owing an event.&lt;/p&gt;

&lt;h2&gt;
  
  
  When a sync enqueue is enough
&lt;/h2&gt;

&lt;p&gt;Skip the outbox when there is no second system. A column update that nothing else consumes is one write. An outbox row would publish a message nobody reads.&lt;/p&gt;

&lt;p&gt;Skip it when you would rather fail the signup than accept a user with no event. Publish to the broker before commit. If that call fails, roll back and return an error. The client retries the whole command. You still need the idempotency key so the retry is not a second user, and a broker outage is an outage of the route. That trade is right when the event is part of success, not a follow-up you can owe.&lt;/p&gt;

&lt;p&gt;Skip it when a miss is repaired without a queue. A search index that rebuilds from the table, or a welcome email the user resends from a button, does not need a durable event. The outbox is for an event you cannot rebuild, or cannot afford to miss until someone notices.&lt;/p&gt;

&lt;p&gt;A publish after commit, with no outbox row, is the dual write again. It is enough only where a miss is repaired by retrying the request, by a rebuild, or by a button. It is not enough for a 201 returned while the broker was down.&lt;/p&gt;

&lt;p&gt;OTF kits do not ship an outbox table or a relay yet. The route you own still commits the business row and the event in one transaction, and a process that stays up still publishes what that transaction left behind.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://microservices.io/patterns/data/transactional-outbox.html" rel="noopener noreferrer"&gt;Transactional outbox&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.postgresql.org/docs/current/sql-select.html#SQL-FOR-UPDATE-SHARE" rel="noopener noreferrer"&gt;SELECT, including FOR UPDATE SKIP LOCKED&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.postgresql.org/docs/current/sql-listen.html" rel="noopener noreferrer"&gt;LISTEN&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.postgresql.org/docs/current/sql-notify.html" rel="noopener noreferrer"&gt;NOTIFY&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/idempotency-keys-owned-api" rel="noopener noreferrer"&gt;Idempotency keys on an owned API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/dead-letter-queue-background-jobs" rel="noopener noreferrer"&gt;Park failed jobs in a DLQ&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/ai-production-background-jobs" rel="noopener noreferrer"&gt;Background jobs for AI features&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>architecture</category>
      <category>backend</category>
      <category>agents</category>
    </item>
    <item>
      <title>An owned API stores idempotency keys in Postgres and replays the first response</title>
      <dc:creator>Dave Kurian</dc:creator>
      <pubDate>Fri, 25 Sep 2026 07:50:11 +0000</pubDate>
      <link>https://dev.to/davekurian/an-owned-api-stores-idempotency-keys-in-postgres-and-replays-the-first-response-55m1</link>
      <guid>https://dev.to/davekurian/an-owned-api-stores-idempotency-keys-in-postgres-and-replays-the-first-response-55m1</guid>
      <description>&lt;p&gt;A retry, a double-click, and a webhook redelivery are the same failure when the handler charges or inserts. The second attempt has to come back with the first result: a key, a row in Postgres, and a rule for the request that is still running.&lt;/p&gt;

&lt;p&gt;The stack is a Next.js route handler in front of Supabase Postgres. The guarantee is a unique constraint and &lt;code&gt;INSERT ... ON CONFLICT&lt;/code&gt;, not a map in the process. &lt;a href="https://otf-kit.dev/blog/blue-green-vs-rolling-owned-backend" rel="noopener noreferrer"&gt;Two colors of a deploy share one database&lt;/a&gt;. A key held in memory splits when the retry lands on the other color. &lt;a href="https://otf-kit.dev/blog/graceful-shutdown-drain-owned-backend" rel="noopener noreferrer"&gt;A drain that ends the process&lt;/a&gt; does the same once that handler is gone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which writes need a key
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.rfc-editor.org/rfc/rfc9110.html#name-idempotent-methods" rel="noopener noreferrer"&gt;RFC 9110&lt;/a&gt; treats GET, HEAD, PUT, DELETE, OPTIONS, and TRACE as idempotent. POST and PATCH are not. Put a key on a call that can charge, book, grant, send, or insert a row the caller would notice twice: create-order, capture-payment, grant-credit, send-invoice, or a webhook with a side effect.&lt;/p&gt;

&lt;p&gt;A read can run twice, and so can a PUT that replaces a resource with a full representation. The command that appends needs a record of the first attempt. &lt;a href="https://otf-kit.dev/blog/circuit-breakers-owned-backend" rel="noopener noreferrer"&gt;A circuit breaker&lt;/a&gt; stops the process waiting on a vendor that is already failing. It does not remember that the write succeeded. The idempotency row does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the key lives
&lt;/h2&gt;

&lt;p&gt;The caller sends &lt;code&gt;Idempotency-Key&lt;/code&gt; when this click is the same attempt. The same SKU and address, five minutes apart, can be a second purchase. Mint the UUID once per button press.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.stripe.com/api/idempotent_requests" rel="noopener noreferrer"&gt;Stripe's API&lt;/a&gt; takes that header on POST, suggests a v4 UUID or another high-entropy string, caps the value at 255 characters, and tells you to keep personal identifiers out of the key. The &lt;a href="https://www.ietf.org/archive/id/draft-ietf-httpapi-idempotency-key-header-07.html" rel="noopener noreferrer"&gt;HTTPAPI Idempotency-Key draft&lt;/a&gt; (draft-07, October 2025) asks for the same uniqueness. It is an Internet-Draft, not an RFC, and the &lt;a href="https://datatracker.ietf.org/doc/draft-ietf-httpapi-idempotency-key-header/" rel="noopener noreferrer"&gt;datatracker entry&lt;/a&gt; lists that revision expired. Copy the behavior. Publish your own retention window. The draft leaves expiry to the resource.&lt;/p&gt;

&lt;p&gt;Verify the webhook signature, then use the id on the verified event. A header the sender can change is a new attempt. Scope a derived id to the endpoint or connected account, and a client key to the session owner. The unique pair is &lt;code&gt;(owner_id, idempotency_key)&lt;/code&gt;. The draft's security notes want the session owner in that pair, so a short key cannot read another caller's stored response.&lt;/p&gt;

&lt;p&gt;A missing key is HTTP 400 on a route that requires the header.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg05f55cpevu9zri4do6c.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg05f55cpevu9zri4do6c.png" alt="Client Idempotency-Key header versus server-derived id into the same key vault" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How long you keep the row
&lt;/h2&gt;

&lt;p&gt;Stripe removes idempotency keys after they are at least 24 hours old. A phone that queues an order for several days needs a longer window. Publish that window next to the header name. After &lt;code&gt;expires_at&lt;/code&gt;, the same key is a new operation.&lt;/p&gt;

&lt;p&gt;Delete expired rows. Index &lt;code&gt;expires_at&lt;/code&gt; so the delete is not a sequential scan. The stored body can hold a name, an address, or a client secret, and past the published window it is not a debugging archive. Recover a crashed charge only while Stripe still remembers the outbound key.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you store with the key
&lt;/h2&gt;

&lt;p&gt;Store the scoped key, a request hash, a state, the response status, the response body, and &lt;code&gt;expires_at&lt;/code&gt;. Hash method, path, and the fields that define the operation, in a fixed order. The draft allows a checksum of the payload or of selected fields. &lt;code&gt;JSON.stringify&lt;/code&gt; of a parsed object, with no stable key order, answers 422 when a retry only reorders fields. Include the path, or a key for &lt;code&gt;POST /api/orders&lt;/code&gt; can replay against &lt;code&gt;POST /api/refunds&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Stripe saves the status and body of the first execution that started, including a 500. Validation failures and in-flight conflicts are not saved. Validate before you claim the key, so a 400 leaves it free. Once the row exists, a different hash is 422 and the second body does not run. Store a failure after an insert or a charge, or the retry repeats that side effect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Concurrent duplicates
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;INSERT ... ON CONFLICT DO NOTHING&lt;/code&gt; is the claim. &lt;a href="https://www.postgresql.org/docs/current/ddl-constraints.html" rel="noopener noreferrer"&gt;Unique constraints&lt;/a&gt; and &lt;a href="https://www.postgresql.org/docs/current/sql-insert.html" rel="noopener noreferrer"&gt;&lt;code&gt;ON CONFLICT&lt;/code&gt;&lt;/a&gt; close the select-then-insert race.&lt;/p&gt;

&lt;p&gt;A double-click does not retry a 409, so the loser waits on &lt;code&gt;SELECT ... FOR UPDATE&lt;/code&gt;. The &lt;a href="https://www.postgresql.org/docs/current/explicit-locking.html" rel="noopener noreferrer"&gt;row lock&lt;/a&gt; holds until the winner commits, then the loser returns the stored status and body. The route timeout caps that wait. A client that retries non-success can take 409 while the first request runs. Stripe saves nothing for that conflict. If the locked read finds no row, the winner rolled back. Return 409.&lt;/p&gt;

&lt;p&gt;Return non-2xx until a webhook stores success, then return the stored 2xx. A 409 after success keeps that vendor going. Commit &lt;code&gt;in_progress&lt;/code&gt; before the charge, call Stripe with the same key or a fixed derivative, and write the response in a second transaction. A new Stripe key on the crash retry is a second charge. The handler below commits the claim, the order, and the 201 in one transaction.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh7wbbjjsigfwb5ble62w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh7wbbjjsigfwb5ble62w.png" alt="Unique constraint claiming one idempotency row while a duplicate request bounces off" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The table and the route
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;owner_id&lt;/code&gt; is the session user, or the endpoint or account for a derived id. &lt;code&gt;response_status&lt;/code&gt; stays null until the handler finishes. supabase-js runs each statement alone, so the claim and the insert share one Postgres transaction, or one SQL function called with RPC. &lt;code&gt;requireUser&lt;/code&gt;, &lt;code&gt;parseOrder&lt;/code&gt;, &lt;code&gt;sha256Hex&lt;/code&gt;, &lt;code&gt;insertOrder&lt;/code&gt;, and &lt;code&gt;withTransaction&lt;/code&gt; belong to the route.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;create&lt;/span&gt; &lt;span class="k"&gt;table&lt;/span&gt; &lt;span class="n"&gt;idempotency_keys&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;owner_id&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;idempotency_key&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;request_hash&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;state&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt; &lt;span class="k"&gt;check&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;state&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'in_progress'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'completed'&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
  &lt;span class="n"&gt;response_status&lt;/span&gt; &lt;span class="nb"&gt;integer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;response_body&lt;/span&gt; &lt;span class="n"&gt;jsonb&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="n"&gt;timestamptz&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
  &lt;span class="n"&gt;expires_at&lt;/span&gt; &lt;span class="n"&gt;timestamptz&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;primary&lt;/span&gt; &lt;span class="k"&gt;key&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;owner_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;idempotency_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;create&lt;/span&gt; &lt;span class="k"&gt;index&lt;/span&gt; &lt;span class="n"&gt;idempotency_keys_expires_at_idx&lt;/span&gt;
  &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="n"&gt;idempotency_keys&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;expires_at&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;POST&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ownerId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;requireUser&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;idempotency-key&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)?.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;invalid_idempotency_key&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;400&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;unknown&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;payload&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;invalid_json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;400&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;order&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;parseOrder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;400&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;requestHash&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sha256Hex&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/api/orders&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;order&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;outcome&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;withTransaction&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;tx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="s2"&gt;`delete from idempotency_keys
        where owner_id = $1 and idempotency_key = $2 and expires_at &amp;lt;= now()`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;ownerId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;claimed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;tx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="s2"&gt;`insert into idempotency_keys
         (owner_id, idempotency_key, request_hash, state, expires_at)
       values ($1, $2, $3, 'in_progress', now() + interval '24 hours')
       on conflict (owner_id, idempotency_key) do nothing
       returning owner_id`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;ownerId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;requestHash&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;claimed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;found&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;tx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s2"&gt;`select request_hash, state, response_status, response_body
           from idempotency_keys
          where owner_id = $1 and idempotency_key = $2
          for update`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;ownerId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
      &lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;found&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;retry&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;request_hash&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="nx"&gt;requestHash&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;mismatch&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;completed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;in_progress&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;replay&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;response_status&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;response_body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;};&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;created&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;insertOrder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ownerId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;created&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;tx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="s2"&gt;`update idempotency_keys
          set state = 'completed', response_status = 201, response_body = $3::jsonb
        where owner_id = $1 and idempotency_key = $2`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;ownerId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;)],&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;created&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;201&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;mismatch&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;idempotency_key_reused&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;422&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;retry&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;in_progress&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;idempotency_in_progress&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;409&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;parseOrder&lt;/code&gt; keeps a stable field order. The delete removes only an expired row. In this shape, &lt;code&gt;for update&lt;/code&gt; sees the winner's committed &lt;code&gt;completed&lt;/code&gt; row.&lt;/p&gt;

&lt;h2&gt;
  
  
  When you don't need this
&lt;/h2&gt;

&lt;p&gt;A GET does not need a key. A PUT of the representation the client already holds does not need one. A DELETE by primary key is finished when the row is absent. Add a key on delete only when the retry must receive the original body.&lt;/p&gt;

&lt;p&gt;A unique email, an overlapping booking, or one ledger line per natural id can already reject the second write. Add this table when the retry must receive the original response, or when the same body is allowed twice and only the attempt id tells them apart. Leave page views, beacons, and intentional counters unkeyed.&lt;/p&gt;

&lt;p&gt;The kits do not ship an idempotency-key layer, so on an API you own this table is the piece you add where a retry can charge or create twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.stripe.com/api/idempotent_requests" rel="noopener noreferrer"&gt;Stripe idempotent requests&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.ietf.org/archive/id/draft-ietf-httpapi-idempotency-key-header-07.html" rel="noopener noreferrer"&gt;Idempotency-Key HTTP header field, draft-ietf-httpapi-idempotency-key-header-07&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://datatracker.ietf.org/doc/draft-ietf-httpapi-idempotency-key-header/" rel="noopener noreferrer"&gt;IETF datatracker entry for the Idempotency-Key draft&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.rfc-editor.org/rfc/rfc9110.html#name-idempotent-methods" rel="noopener noreferrer"&gt;RFC 9110, idempotent methods&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.postgresql.org/docs/current/sql-insert.html" rel="noopener noreferrer"&gt;PostgreSQL INSERT, ON CONFLICT&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.postgresql.org/docs/current/ddl-constraints.html" rel="noopener noreferrer"&gt;PostgreSQL unique constraints&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.postgresql.org/docs/current/explicit-locking.html" rel="noopener noreferrer"&gt;PostgreSQL row-level locks&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/blue-green-vs-rolling-owned-backend" rel="noopener noreferrer"&gt;Blue-green versus rolling on an owned backend&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/graceful-shutdown-drain-owned-backend" rel="noopener noreferrer"&gt;Graceful shutdown and drain on an owned backend&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/circuit-breakers-owned-backend" rel="noopener noreferrer"&gt;Circuit breakers on an owned backend&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>architecture</category>
      <category>backend</category>
      <category>agents</category>
    </item>
    <item>
      <title>When to cut traffic to a warm standby instead of a rolling replace</title>
      <dc:creator>Dave Kurian</dc:creator>
      <pubDate>Fri, 25 Sep 2026 05:31:12 +0000</pubDate>
      <link>https://dev.to/davekurian/when-to-cut-traffic-to-a-warm-standby-instead-of-a-rolling-replace-d3g</link>
      <guid>https://dev.to/davekurian/when-to-cut-traffic-to-a-warm-standby-instead-of-a-rolling-replace-d3g</guid>
      <description>&lt;p&gt;&lt;a href="https://otf-kit.dev/blog/rolling-deploys-owned-backend" rel="noopener noreferrer"&gt;Rolling deploys&lt;/a&gt; are the default on an owned backend: one fleet, old and new serving together until the old instances are gone. Blue-green is the other move. A second stack stays warm, and you cut traffic to it in one switch.&lt;/p&gt;

&lt;p&gt;The choice is whether the two builds can share a request path, how fast you must leave a bad build, and whether the last good state still exists after the new code has written data. A second full stack idles until the cut. Idle memory is a bill you take on purpose.&lt;/p&gt;

&lt;p&gt;Surge, unavailability, and instance order belong to the rolling deploys piece. This post decides when you stop mixing versions and move a load balancer, a Service selector, or a weighted route.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cut when versions must not share a request path
&lt;/h2&gt;

&lt;p&gt;A mixed fleet is safe only while a request can land on either version and finish correctly.&lt;/p&gt;

&lt;p&gt;A wire change breaks that. The new process speaks a shape the old one rejects: a required field, a renamed header, a content type, an RPC the old binary does not register. Callers split across the two shapes. A retry draws the same mix. Failures follow the ratio.&lt;/p&gt;

&lt;p&gt;A schema the old process rejects fails the same way in stored data. Green writes a column, an enum, or a payload blue cannot read. Traffic left on blue errors, including reads of rows green already wrote.&lt;/p&gt;

&lt;p&gt;The write path is stricter. When every peer must already speak the new contract — a lock format, an outbox event, a cache key — the mixed fleet splits writes. Soak green on real dependencies with user traffic held off: same database, same queues, same downstreams. Move the balancer or Service selector once, and green becomes the only version taking production requests.&lt;/p&gt;

&lt;p&gt;Rollback is that switch in reverse. It works while blue is up and can still serve the data that now exists. Keep blue warm until you trust the cut.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cut when time-to-rollback matters more than spare capacity
&lt;/h2&gt;

&lt;p&gt;A gradual replace unwinds through startup, readiness, and the surge budget. That wait fits a slow failure: latency on a slice, an error rate that climbs with the mix. It fits an immediate failure poorly. A bad config, an auth check that rejects every token, or a startup path that corrupts a shared cache should be left before the next health interval.&lt;/p&gt;

&lt;p&gt;A flip moves a selector or a weight from 100/0 to 0/100. The bill is two full stacks, both ready, for as long as the old color stays. On a small service the second Deployment is cheap beside the time you get back. On a memory-heavy tier the idle half is the decision. The same heap doubles resident memory until blue scales down.&lt;/p&gt;

&lt;p&gt;A ramp, small weight then larger, stays gradual. Use it when versions coexist and the blast radius should stay limited. Blue-green here means the full cut onto a stack that was already warm.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqh67mojah6gjg9nfgivj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqh67mojah6gjg9nfgivj.png" alt="Blue-green cutover switch versus rolling batch replace side by side" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Stay gradual when a flip would not restore the last good state
&lt;/h2&gt;

&lt;p&gt;Two colors, one database. Drop a column, rewrite data, or change what a field means, and blue cannot read what green wrote. The selector can point at blue again. Blue then fails on current rows. Sequence the schema as expand, deploy, contract. Expand keeps the old shape valid for every process still running. Deploy writes both shapes, or writes the new one so old code ignores it. Contract removes the old shape only after nothing that might still run depends on it. The standby hands you the previous binary, which has to read the rows green left behind.&lt;/p&gt;

&lt;p&gt;Call green warm only once real requests can succeed. Pools are open. Hot-path caches are filled, or misses are priced and the downstream is sized. A shadow load has run that path against production dependencies, users still on blue. A readiness route that returns 200 before any real query will take the herd. Readiness is the ability to do the work. A first wave that opens pools and times out downstreams survives one probe, then falls over.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://otf-kit.dev/blog/graceful-shutdown-drain-owned-backend" rel="noopener noreferrer"&gt;Graceful shutdown and drain&lt;/a&gt; still apply on the way out. In-flight work on blue continues after the selector moves. Drain it. New requests go to green. Old ones finish, or hit their deadline, on blue. Scale blue down after the drain, once you are finished with the reverse switch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stay gradual when versions can share the path
&lt;/h2&gt;

&lt;p&gt;Stay on a rolling replace when the versions can coexist. A parser, a timeout, a log field, a query that returns the same rows: one contract, and a retry may land on either build. This change ships on the fleet you already have.&lt;/p&gt;

&lt;p&gt;Stay there when the copy will not fit. A cache sized to the host, a heap tuned to the box, or one GPU means the overlap is another tier. Rolling spends the surge already configured, a slice of the same fleet.&lt;/p&gt;

&lt;p&gt;Stay there when the fault should hit a slice. A slower query or a higher error rate shows while most instances still run the old build. Stopping the replace leaves most callers on the last good build. A full cut reports the same bug after every caller has moved. Prefer the fast reverse across all traffic, or the slower slice with most of the fleet still on the old build.&lt;/p&gt;

&lt;p&gt;A slice can still stampede a cache, storm a lock, or retry into a downstream. &lt;a href="https://otf-kit.dev/blog/circuit-breakers-owned-backend" rel="noopener noreferrer"&gt;Circuit breakers&lt;/a&gt; cap the amplification while you halt. Whether the two versions can share a path stays a separate decision.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Blue-green cutover&lt;/th&gt;
&lt;th&gt;Rolling gradual replace&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Traffic&lt;/td&gt;
&lt;td&gt;One switch of all new requests onto a warm stack&lt;/td&gt;
&lt;td&gt;Batches, with old and new serving together&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Capacity&lt;/td&gt;
&lt;td&gt;Two full stacks for the overlap&lt;/td&gt;
&lt;td&gt;One fleet, within maxUnavailable and maxSurge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fits when&lt;/td&gt;
&lt;td&gt;A shared path would fail, or rollback must beat the next health interval&lt;/td&gt;
&lt;td&gt;Coexistence is safe, spare capacity is a slice, or the fault should stay small&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rollback&lt;/td&gt;
&lt;td&gt;Reverse the selector or weight while the previous color is up and can read current data&lt;/td&gt;
&lt;td&gt;Stop the replace; the wait is startup and readiness&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data&lt;/td&gt;
&lt;td&gt;The flip leaves migrations and other writes where they are&lt;/td&gt;
&lt;td&gt;Mixed versions split writes when the contract has changed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;You operate&lt;/td&gt;
&lt;td&gt;Two Deployments, a selector or ingress weight, then a drain&lt;/td&gt;
&lt;td&gt;One Deployment; the controller walks the replace&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fji7zo6lp4sbca5hy8kux.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fji7zo6lp4sbca5hy8kux.png" alt="Traffic flip to green warm standby with blue drain for rollback" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The flip puts one version on the request path. The batch replace stays mixed until the old instances are gone.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you move, and what stays up
&lt;/h2&gt;

&lt;p&gt;Kubernetes Deployments perform the gradual replace with maxUnavailable and maxSurge on one Deployment. The cut is two Deployments and a selector or ingress weight you move once green is warm. Warm means pools open, caches filled or misses priced, a shadow load finished, and readiness held closed until then. Drain blue before scaling it to zero. In-flight requests outlive the selector change, and the time blue stays up is the rollback window.&lt;/p&gt;

&lt;p&gt;Where coexistence is safe and the idle cost is already accepted, keep the gradual replace. Where a shared path would fail, or leaving a bad build cannot wait on startup and readiness, cut to the warm standby and keep the previous color until that cut has earned it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/rolling-deploys-owned-backend" rel="noopener noreferrer"&gt;Rolling deploys on an owned backend&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/graceful-shutdown-drain-owned-backend" rel="noopener noreferrer"&gt;Graceful shutdown and drain on an owned backend&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/circuit-breakers-owned-backend" rel="noopener noreferrer"&gt;Circuit breakers on an owned backend&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://kubernetes.io/docs/concepts/workloads/controllers/deployment/" rel="noopener noreferrer"&gt;Kubernetes Deployments&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/whitepapers/latest/blue-green-deployments/welcome.html" rel="noopener noreferrer"&gt;Blue/Green Deployments on AWS&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>architecture</category>
      <category>backend</category>
      <category>agents</category>
    </item>
    <item>
      <title>otf-kit vs Shadcn Studio: a decision checklist before you buy</title>
      <dc:creator>Dave Kurian</dc:creator>
      <pubDate>Fri, 25 Sep 2026 03:15:08 +0000</pubDate>
      <link>https://dev.to/davekurian/otf-kit-vs-shadcn-studio-a-decision-checklist-before-you-buy-400a</link>
      <guid>https://dev.to/davekurian/otf-kit-vs-shadcn-studio-a-decision-checklist-before-you-buy-400a</guid>
      <description>&lt;p&gt;Write down the file you expect on disk tomorrow morning. If the sentence is "a block inside the React app we already run," you are looking at one product. If the sentence is "a repo that already signs a user in and charges a card, on web and on mobile," you are looking at the other. Pay only after that sentence is boring and specific. This checklist is adapted from the full comparison on &lt;a href="https://otf-kit.dev/compare/otf-kit-vs-shadcn-studio" rel="noopener noreferrer"&gt;OTF&lt;/a&gt;, which I update when a price changes.&lt;/p&gt;

&lt;p&gt;I am Dave Kurian, and I build OTF. One branch of this checklist ends in our kit. Another ends in Shadcn Studio. A third ends in spending nothing. I would rather you take the third than buy a skeleton for a button.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sort the job into one sentence
&lt;/h2&gt;

&lt;p&gt;Use one of these, and drop the other:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The app already exists. The missing work is blocks, pages, and templates copied into it, plus a Figma file and a theme tool.&lt;/li&gt;
&lt;li&gt;The product is still a folder. The missing work is login, data, and Stripe, for web and mobile, in a repo you will keep.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both can be true across a year. They are rarely both true this month. Buy the month you are in. Screens: catalog. A stranger who still cannot pay you: kit.&lt;/p&gt;

&lt;p&gt;Shadcn Studio is the catalog, on a one-time license. OTF is a free SDK plus paid kits. The SDK is $0 and MIT. A landing template is a separate $9 file. A kit is the wired product.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checklist A — buy the catalog
&lt;/h2&gt;

&lt;p&gt;Mark yes or no. Buy Shadcn Studio when the yeses are the ones that match your week.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A React app already exists. Accounts and data are already settled.&lt;/li&gt;
&lt;li&gt;You want a dense gallery in that app: marketing blocks, dashboard and application blocks, ecommerce blocks, pages, templates.&lt;/li&gt;
&lt;li&gt;You want the Figma kit, the theme generator, the MCP server, the builder, and the IDE extension on the same license.&lt;/li&gt;
&lt;li&gt;The path you will actually use is the CLI, the MCP server, or a prompt. The work is paste-into-this-repo.&lt;/li&gt;
&lt;li&gt;You have read two clocks, not one: lifetime access, with new components, blocks, and templates monthly, and one year of premium support.&lt;/li&gt;
&lt;li&gt;Day one does not need to charge a card. Your app already knows how, or that work is out of scope.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Smaller door: Flow, one template, was $59 on the 2026-09-24 fetch. The same view also showed "Get all access." If the template is the whole purchase, confirm the control you click is the $59 template.&lt;/p&gt;

&lt;p&gt;Sampling is free. Community is the free tier. The license page adds a Commons Clause to MIT: no resale of the components, and no product that competes with Shadcn Studio. Free MCP reaches free components and blocks only. Paid blocks are a separate gate.&lt;/p&gt;

&lt;p&gt;A kit does not satisfy this list. You would be cloning a login flow to get a dashboard shell.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checklist B — buy the kit
&lt;/h2&gt;

&lt;p&gt;Buy an OTF kit when these are the yeses.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Someone still has to stand up login, stored data, and a card charge.&lt;/li&gt;
&lt;li&gt;Web and mobile should come from one SDK, in one repo you keep.&lt;/li&gt;
&lt;li&gt;You can live with 12 months of updates on a $99 kit, or you want lifetime updates and will take the Everything Bundle or Team.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three kits are on sale at $99: SaaS Dashboard, Fitness, and Arcade. Choose by the product, not by collecting all three. For a standard SaaS, start with SaaS Dashboard, on &lt;a href="https://otf-kit.dev/templates/saas-dashboard" rel="noopener noreferrer"&gt;its template page&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The SDK stays free. There is no paid badge on a component. Checklist B fails, correctly, when the only hole is a control you could ship this afternoon.&lt;/p&gt;

&lt;p&gt;A landing is not a kit. Fifteen landing templates are $9 each. Buy one for a page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checklist C — spend nothing
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;You need a button, a form, or one settings screen. Use the free SDK, or the free Community pieces if the catalog's app is the one you already have.&lt;/li&gt;
&lt;li&gt;You need Booking or Marketplace shipped today. Skip us for now. Booking is a Preview. Marketplace isn't for sale yet and isn't part of the Bundle.&lt;/li&gt;
&lt;li&gt;The real wish was a Figma kit and a monthly run of new blocks for an app that already runs. That wish is checklist A.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Prices to paste into the note
&lt;/h2&gt;

&lt;p&gt;Shadcn Studio numbers are as fetched on 2026-09-24. Check their pricing page before you buy. Struck figures sat beside the live price that day. I am not translating them into a discount. OTF numbers are the public ladder. Team has no printed dollar amount, so this note has none.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Line&lt;/th&gt;
&lt;th&gt;Shadcn Studio, as fetched 2026-09-24&lt;/th&gt;
&lt;th&gt;OTF&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;td&gt;Community, $0&lt;/td&gt;
&lt;td&gt;SDK, $0, MIT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Entry paid&lt;/td&gt;
&lt;td&gt;Basic $99, a $219 figure beside it, 1 seat, 48 business hours&lt;/td&gt;
&lt;td&gt;Landing template $9, fifteen of them&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Main paid&lt;/td&gt;
&lt;td&gt;Pro $199, a $359 figure beside it, marked Best Value, 1 seat&lt;/td&gt;
&lt;td&gt;Kit $99, 12 months of updates: SaaS Dashboard, Fitness, Arcade&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Team&lt;/td&gt;
&lt;td&gt;$449, a $719 figure beside it, 15 seats, 24 business hours&lt;/td&gt;
&lt;td&gt;5 seats, by email, lifetime updates, no dollar amount&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Top&lt;/td&gt;
&lt;td&gt;Enterprise $849, a $1,299 figure beside it, unlimited seats&lt;/td&gt;
&lt;td&gt;Everything Bundle $149, struck $531, Save $382, lifetime updates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single file&lt;/td&gt;
&lt;td&gt;Flow, $59&lt;/td&gt;
&lt;td&gt;A landing, $9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Updates&lt;/td&gt;
&lt;td&gt;One-time payment, lifetime access, new pieces monthly, unlimited projects&lt;/td&gt;
&lt;td&gt;12 months on a kit. Lifetime on the Bundle and on Team&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Support&lt;/td&gt;
&lt;td&gt;FAQ: one year of premium support&lt;/td&gt;
&lt;td&gt;The printed promise is the update window. No support-year figure is added here&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Copy two Studio sentences into the note, separately. You pay once and the FAQ describes later components, blocks, templates, and AI tools as included. Support on that purchase lasts one year. People keep the first sentence and lose the second.&lt;/p&gt;

&lt;p&gt;On our side the mix-up runs the other way. A $99 kit includes 12 months, then you still hold the code, and you do not hold a claim on the following year. The $149 Bundle is the lifetime-updates line. Team updates are lifetime too.&lt;/p&gt;

&lt;p&gt;OTF figures live on &lt;a href="https://otf-kit.dev/pricing" rel="noopener noreferrer"&gt;the pricing page&lt;/a&gt;. Kit status changes are posted on &lt;a href="https://otf-kit.dev/changelog" rel="noopener noreferrer"&gt;the changelog&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Score the sheet once
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Checklist A matches the week: buy Shadcn Studio, or Flow alone if one template is the deliverable.&lt;/li&gt;
&lt;li&gt;Checklist B matches the week: buy one named kit, or the Bundle if you want lifetime updates, the live kits, and the fifteen landings.&lt;/li&gt;
&lt;li&gt;Checklist C matches: do not pay. Stay on the free SDK.&lt;/li&gt;
&lt;li&gt;A and B both look half true: buy the artifact for this month's work. A catalog license leaves payments in the app you already wrote. A bundle does not include their Figma kit, their theme generator, or their block gallery.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I keep the full, updated comparison on &lt;a href="https://otf-kit.dev/compare/otf-kit-vs-shadcn-studio" rel="noopener noreferrer"&gt;OTF&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>react</category>
      <category>saas</category>
      <category>reactnative</category>
    </item>
    <item>
      <title>Cursor Automations last mile: Security Review plus Rollouts, not hope-CI</title>
      <dc:creator>Dave Kurian</dc:creator>
      <pubDate>Thu, 24 Sep 2026 11:39:50 +0000</pubDate>
      <link>https://dev.to/davekurian/cursor-automations-last-mile-security-review-plus-rollouts-not-hope-ci-4fgc</link>
      <guid>https://dev.to/davekurian/cursor-automations-last-mile-security-review-plus-rollouts-not-hope-ci-4fgc</guid>
      <description>&lt;p&gt;Last-mile confidence is not a vibe and it is not “CI was green.” For Teams and Enterprise, Cursor Automations now cover the two gaps that still eat merge-and-ship weekends: exploitable-bug review on pull requests, and per-environment deploy health after merge—with notify and a revert path when a named regression shows up. Bugbot stays the style and nits bot. Do not replace it. Do not treat Expo OTA staged updates as the same product.&lt;/p&gt;

&lt;p&gt;The buyer question is specific: how do you enable Cursor Automations last-mile bots so every ready PR gets an attack-path-plus-fix security pass before merge, and so each environment gets deploy-health and named-regression handling after—without throwing out Bugbot? One stack answers it. Security Review on the PR. Rollouts on the deploy. Same Automations surface. One last mile.&lt;/p&gt;

&lt;p&gt;This is not two co-equal posts. Security Review without Rollouts is a merge ritual that still hopes production is fine. Rollouts without Security Review is a health dashboard on a change nobody attacked on the way in. Together they are last-mile confidence. &lt;a href="https://otf-kit.dev/blog/eas-update-staged-rollout-guide" rel="noopener noreferrer"&gt;Expo’s staged EAS Update guide&lt;/a&gt; is a different last mile: JS bundle percentage on devices. Cursor Rollouts are PR-to-environment health, notify, and revert for the thing you just merged. Do not mix the runbooks.&lt;/p&gt;

&lt;p&gt;Primary source: Cursor’s changelog for Rollouts and the security reviewer (2026-09-23) at &lt;a href="https://cursor.com/changelog/rollouts-and-security-reviewer" rel="noopener noreferrer"&gt;cursor.com/changelog/rollouts-and-security-reviewer&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What last-mile actually means here
&lt;/h2&gt;

&lt;p&gt;Most teams already have required checks, CODEOWNERS, and a bot that comments on unused imports. That is not last mile. Last mile is the window where a ready PR can still ship an exploitable bug, and the window after merge where a “healthy” deploy is actually a named regression in one environment.&lt;/p&gt;

&lt;p&gt;Security Review is the first window: exploitable issues with an attack path a reviewer can follow and a fix path a patcher can take. It is not a style pass. Rollouts is the second window: after merge, health per named environment, notify when it is not, and a revert path when a named regression is the story. Bugbot stays the third bot—style and consistency—cheap on purpose. Folding style into Security Review trains people to skip security comments. Folding security into Bugbot trains people to treat attack paths as nits. Keep the lanes. Pair this with an &lt;a href="https://otf-kit.dev/blog/ai-app-security-checklist" rel="noopener noreferrer"&gt;AI app security checklist&lt;/a&gt; for boundaries Security Review cannot invent for you.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4oyxtlhf2lr26e04gl62.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4oyxtlhf2lr26e04gl62.png" alt="Hope CI was enough versus Security Review plus Rollouts Automations" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Hope-CI-was-enough looks like: tests passed, linter passed, two approvals, merge, assume staging and prod will tell you later. Security Review + Rollouts Automations looks like: every ready PR gets exploitable-bug review with attack path and fix; after merge, Rollouts watches per-environment health, notifies on named regressions, and can open a revert PR for review. Same PR object. Different failure modes. Different bots.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enable Security Review on every ready PR
&lt;/h2&gt;

&lt;p&gt;Treat ready (non-draft) as the trigger, not “when someone remembers to @ the bot.” Drafts are for incomplete work. The moment a PR leaves draft, it is in the merge funnel.&lt;/p&gt;

&lt;p&gt;Wire the Automation so:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Scope is Teams/Enterprise Automations, not a personal rule you cannot audit.&lt;/li&gt;
&lt;li&gt;Event is ready pull request (opened ready, or converted from draft). Re-runs on relevant pushes keep the review honest as the diff moves.&lt;/li&gt;
&lt;li&gt;Output is exploitable-bug review: attack path plus fix. If the model cannot name a path, it should not invent a CVE-shaped essay.&lt;/li&gt;
&lt;li&gt;Bugbot is unchanged. Do not disable Bugbot to “reduce bots.” Reduce overlap by job, not by count.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What “attack path + fix” means in practice: untrusted input reaches a sink (authz skip, template, shell, deserialization, IDOR), and the fix is the concrete change in this PR—guard, query bound, deny-by-default, or drop the endpoint. If the Automation cannot point at a hunk, it is not last-mile.&lt;/p&gt;

&lt;p&gt;Required status is a policy choice. If Security Review is advisory, people will merge around it the week you are busy. Teams that already require Bugbot for nits should not assume that requirement covers exploitable bugs. It does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enable Rollouts for PR → environment health
&lt;/h2&gt;

&lt;p&gt;After merge, the PR is a candidate in each environment you actually ship to. Rollouts attaches monitoring to the change, reports health per environment, notifies when health is not the story you wanted, and—depending on configuration—can open a revert PR for review or hand the finding to a cloud agent. Per Cursor’s changelog, it does not merge or roll back on its own today.&lt;/p&gt;

&lt;p&gt;Name environments the way you operate (&lt;code&gt;preview&lt;/code&gt;, &lt;code&gt;staging&lt;/code&gt;, &lt;code&gt;prod&lt;/code&gt;). Rollouts is not Expo’s rollout percent. For OTA percentages, stay on the &lt;a href="https://otf-kit.dev/blog/eas-update-staged-rollout-guide" rel="noopener noreferrer"&gt;EAS Update staged rollout guide&lt;/a&gt;. For “this merged PR’s deploy is sick in staging,” stay on Cursor Rollouts.&lt;/p&gt;

&lt;p&gt;Health has to be named. Named regression means a check, error class, or SLO you already believe—checkout, auth, a critical job, a latency budget you page on. Wire notify to the people who can revert. You should answer: which PR, which env, which signal, who was notified, was a revert opened.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk75pa482m392s0v2i2d9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk75pa482m392s0v2i2d9.png" alt="Ready PR to Security Review to merge to Rollouts per-env health to notify or revert" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Nothing in that spine is Bugbot. Nothing is an Expo runtime version. Ready is the security gate. Merge is the handoff. Per-env health is the deploy gate. Notify and revert are how you leave without a war room as the default.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do not confuse Cursor Rollouts with Expo staged OTA
&lt;/h2&gt;

&lt;p&gt;Expo / EAS Update staged rollout ships a JavaScript bundle to a fraction of clients and watches crash and runtime signals. The unit is the update on a channel. The client is a device. Use &lt;a href="https://otf-kit.dev/blog/eas-update-staged-rollout-guide" rel="noopener noreferrer"&gt;eas-update-staged-rollout-guide&lt;/a&gt; when the question is “what fraction of phones get this bundle.”&lt;/p&gt;

&lt;p&gt;Cursor Rollouts (Automations, Teams/Enterprise, changelog 2026-09-23) asks whether the merged PR’s deploy is healthy in each environment, who hears about a named regression, and whether a revert path opens. The blast radius is an environment, not a device cohort. You can run both in one company. Document them as two last miles.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep Bugbot for style
&lt;/h2&gt;

&lt;p&gt;Bugbot owns dead code, naming, and review hygiene. Security Review owns exploitable bugs on ready PRs. Rollouts owns post-merge per-environment health, notify, and revert path. Three bots is not a failure. One bot with three jobs is. For owned backends you already operate, keep &lt;a href="https://otf-kit.dev/blog/production-structured-logging-for-agents" rel="noopener noreferrer"&gt;structured logging agents can triage&lt;/a&gt; so Rollouts has signals worth trusting.&lt;/p&gt;

&lt;h2&gt;
  
  
  A enablement path that does not stall
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Confirm Automations on Teams or Enterprise.&lt;/li&gt;
&lt;li&gt;Turn on Security Review for ready PRs. Convert-from-draft must count. Sample a week of comments: keep only findings with a path and a fix.&lt;/li&gt;
&lt;li&gt;Leave Bugbot on. If volume is high, tune Bugbot, not Security Review.&lt;/li&gt;
&lt;li&gt;Turn on Rollouts for the environments you deploy. Attach signals you already trust. Wire notify to the on-call who can revert.&lt;/li&gt;
&lt;li&gt;Practice one revert in non-prod so revert is not a rumor.&lt;/li&gt;
&lt;li&gt;Write two sentences in the team doc: Cursor Rollouts ≠ EAS staged updates; Bugbot ≠ Security Review.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The failure mode is a fourth bot that restates CI. The other failure mode is skipping Security Review because Rollouts “will catch it.” Rollouts catch deploy health. They do not reconstruct an IDOR from a 200 OK.&lt;/p&gt;

&lt;h2&gt;
  
  
  What “good” looks like after a week
&lt;/h2&gt;

&lt;p&gt;A ready PR that touches auth gets a Security Review comment that names the path and the fix. Bugbot still nags about a leftover &lt;code&gt;console.log&lt;/code&gt;. A merge to staging trips a named regression; Rollouts notifies; a revert path opens in the same spine. If Security Review only ever says “looks fine” on risky diffs, tighten the ready-PR trigger and demand path + fix. If someone asks whether this replaces Expo rollout percent, send them the EAS guide and this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we are not claiming
&lt;/h2&gt;

&lt;p&gt;The changelog is the source for Rollouts and the security reviewer as of 2026-09-23. This post does not invent plan SKUs beyond Teams/Enterprise Automations, does not invent package versions, and does not quote unpublished metrics. It does not claim Security Review finds every exploitable bug. It claims you can put exploitable-bug review with attack path and fix on ready PRs, keep Bugbot on style, and put PR-to-environment health, notify, and a revert path on the other side of merge.&lt;/p&gt;

&lt;p&gt;Last-mile confidence is that pair, on Automations, not a feeling that CI was enough.&lt;/p&gt;

&lt;p&gt;Owned kits still help when the app itself is the thing you review and roll out: &lt;a href="https://otf-kit.dev/templates" rel="noopener noreferrer"&gt;browse templates&lt;/a&gt; for full-stack starting points with agent configs already in the repo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://cursor.com/changelog/rollouts-and-security-reviewer" rel="noopener noreferrer"&gt;Cursor changelog: Rollouts and security reviewer (2026-09-23)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/eas-update-staged-rollout-guide" rel="noopener noreferrer"&gt;EAS Update staged rollout guide (Expo OTA — different last mile)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/ai-app-security-checklist" rel="noopener noreferrer"&gt;AI app security checklist&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/production-structured-logging-for-agents" rel="noopener noreferrer"&gt;Structured production logs with correlation IDs agents can triage&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>cursor</category>
      <category>agents</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Wire TanStack AI agents to OAuth-protected MCP via Vercel Connect</title>
      <dc:creator>Dave Kurian</dc:creator>
      <pubDate>Thu, 24 Sep 2026 09:45:09 +0000</pubDate>
      <link>https://dev.to/davekurian/wire-tanstack-ai-agents-to-oauth-protected-mcp-via-vercel-connect-56o3</link>
      <guid>https://dev.to/davekurian/wire-tanstack-ai-agents-to-oauth-protected-mcp-via-vercel-connect-56o3</guid>
      <description>&lt;p&gt;If you already run TanStack AI agents that call MCP tools, the boring failure mode is not that the model is dumb. It is that you pasted a long-lived MCP token into env, shipped it with the deployment, and now every rotate is a redeploy plus a hope that no screenshot of &lt;code&gt;$CONNECT_*&lt;/code&gt; ever leaked. Vercel’s changelog for &lt;a href="https://vercel.com/changelog/vercel-connect-tanstack-ai" rel="noopener noreferrer"&gt;Vercel Connect + TanStack AI&lt;/a&gt; (2026-09-24) is aimed at that boundary: install &lt;code&gt;@vercel/connect/tanstack-ai&lt;/code&gt;, send MCP traffic through &lt;code&gt;connectMCPTransport&lt;/code&gt;, and when the server returns a consent challenge, catch &lt;code&gt;getConsentChallenge&lt;/code&gt; and redirect the user. Tokens are issued at runtime at an owned route. They do not live in env for you to rotate.&lt;/p&gt;

&lt;p&gt;This post is the keep-path PSEO for that wiring. It is not the EAS/Claude connector path in &lt;a href="https://otf-kit.dev/blog/expo-mcp-connector-claude" rel="noopener noreferrer"&gt;Expo MCP connector for Claude&lt;/a&gt;. Different runtime, different consent surface, different place the secret is allowed to exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Connect is doing at the agent boundary
&lt;/h2&gt;

&lt;p&gt;TanStack AI already knows how to call tools. MCP already knows how to expose tools behind OAuth. The missing piece on a Vercel deployment is a first-party transport that refuses to treat “bearer in &lt;code&gt;$CONNECT_TOKEN&lt;/code&gt;” as the integration.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;connectMCPTransport&lt;/code&gt; is that transport. You point the agent at MCP the same way you would any other tool host, except the bytes do not leave your app with a static secret. When the MCP server needs a user (or org) to grant access, Connect surfaces a consent challenge. Your route handles it with &lt;code&gt;getConsentChallenge&lt;/code&gt; and redirects. After consent, a runtime token is bound to that session at the route you own.&lt;/p&gt;

&lt;p&gt;That is the product claim worth keeping: &lt;strong&gt;owned-route MCP auth&lt;/strong&gt;. The agent never becomes a secret store. Env never becomes a token locker. Rotation is “consent again,” not “edit production env and pray.”&lt;/p&gt;

&lt;p&gt;If you still think in gateway terms — one billed path for model calls, one for tools — pair this with &lt;a href="https://otf-kit.dev/blog/gpt-live-1-delegation-ai-gateway" rel="noopener noreferrer"&gt;AI Gateway delegation&lt;/a&gt;. Gateway is the model hop. Connect is the OAuth hop for MCP. Do not collapse them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Install and the only import that matters
&lt;/h2&gt;

&lt;p&gt;From the changelog: install &lt;code&gt;@vercel/connect/tanstack-ai&lt;/code&gt;. That package is the TanStack-specific adapter. Do not cargo-cult a mobile EAS connector, a generic MCP SDK auth helper, or a “put the token in &lt;code&gt;$CONNECT_MCP_TOKEN&lt;/code&gt;” snippet from an older thread.&lt;/p&gt;

&lt;p&gt;Keep Connect configuration in env as &lt;em&gt;connection&lt;/em&gt; config — whatever &lt;code&gt;$CONNECT_*&lt;/code&gt; names your dashboard emits — not as the user’s MCP access token. Path-only on the app side: your route owns consent and the agent endpoint.&lt;/p&gt;

&lt;p&gt;Honest skeleton from the changelog: TanStack AI (&lt;code&gt;chat&lt;/code&gt;, &lt;code&gt;createMCPClient&lt;/code&gt;) with MCP attached through &lt;code&gt;connectMCPTransport&lt;/code&gt; — not a raw fetch with &lt;code&gt;Authorization&lt;/code&gt; from &lt;code&gt;.env&lt;/code&gt;. The same HTTP handler catches &lt;code&gt;getConsentChallenge&lt;/code&gt; and redirects; it does not retry the model with a guessed token. Storing the resulting token in env to “make CI simpler” undoes the feature.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;connectMCPTransport&lt;/code&gt;: keep the agent ignorant of OAuth
&lt;/h2&gt;

&lt;p&gt;The agent’s job is tool names, arguments, and when to stop. OAuth is not a tool. If you shove authorize URLs into system prompts, you will get a model that pastes tokens into logs.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;connectMCPTransport&lt;/code&gt; sits under the tool layer. The provider is called before every MCP request, so the token is always fresh. Failed auth is not a chat message; it is a control-flow exception your route understands. That is why &lt;code&gt;getConsentChallenge&lt;/code&gt; exists as a catchable object rather than a string the model might echo.&lt;/p&gt;

&lt;p&gt;Trace transport errors and redirects, not bearer prefixes. Treat consent timeouts as user-wait. Multiple MCP servers can share the Connect consent pattern — not one shared env token.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;getConsentChallenge&lt;/code&gt; and the redirect you actually ship
&lt;/h2&gt;

&lt;p&gt;Consent is a browser problem. When &lt;code&gt;getConsentChallenge&lt;/code&gt; fires on the agent route the UI posts to, that route should:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Stop the current agent turn. Do not half-apply tool results.&lt;/li&gt;
&lt;li&gt;Redirect the user to the consent URL Connect gave you (changelog uses a 303).&lt;/li&gt;
&lt;li&gt;Land back on a path you own after the user accepts or denies.&lt;/li&gt;
&lt;li&gt;Resume the agent only after a runtime token exists for that session.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;“Owned route” is load-bearing. If consent returns to a third-party page you do not control, someone else’s cookie becomes your security model. Keep the callback on your deployment — path-only, like &lt;code&gt;/connect/consent&lt;/code&gt; next to the agent route.&lt;/p&gt;

&lt;p&gt;Denied consent is first-class. Show that MCP tools are unavailable. Do not fall back to a pasted env token “just this once.”&lt;/p&gt;

&lt;p&gt;Catch the challenge &lt;em&gt;before&lt;/em&gt; the model runs. The changelog is explicit: if a consent error is raised inside a tool call instead, it reaches the model as an error string rather than the user as a redirect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Runtime tokens vs env tokens
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F60t1f4286xqfzatlzbbw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F60t1f4286xqfzatlzbbw.png" alt="Pasted env MCP token vs Connect consent at owned route" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Pasted MCP token in env&lt;/th&gt;
&lt;th&gt;Connect consent at the route&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Where the secret lives&lt;/td&gt;
&lt;td&gt;Deployment env, often copied into previews&lt;/td&gt;
&lt;td&gt;Issued at runtime after &lt;code&gt;getConsentChallenge&lt;/code&gt; redirect&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rotate&lt;/td&gt;
&lt;td&gt;Change env, redeploy, invalidate every replica&lt;/td&gt;
&lt;td&gt;Re-consent; old runtime token dies with the session&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blast radius&lt;/td&gt;
&lt;td&gt;Anyone with env or a leaked preview&lt;/td&gt;
&lt;td&gt;Bound to the user/session that consented&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent code&lt;/td&gt;
&lt;td&gt;Tempted to log headers&lt;/td&gt;
&lt;td&gt;Transport handles OAuth; agent sees tools&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fits TanStack on Vercel&lt;/td&gt;
&lt;td&gt;Works until the first leak&lt;/td&gt;
&lt;td&gt;What &lt;code&gt;@vercel/connect/tanstack-ai&lt;/code&gt; is for&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Env is fine for &lt;em&gt;which&lt;/em&gt; Connect app you are, &lt;code&gt;$CONNECT_*&lt;/code&gt; identifiers, and public MCP URLs. Env is not fine for the OAuth access token that talks to a customer’s MCP. If your runbook still says “rotate MCP_TOKEN quarterly,” you adopted a rename, not Connect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sequence you can keep in your head
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fst5q9okws2qvapog5rdl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fst5q9okws2qvapog5rdl.png" alt="TanStack agent to connectMCPTransport to consent to runtime token to OAuth MCP" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;UI hits your owned agent route.&lt;/li&gt;
&lt;li&gt;TanStack AI starts a turn and needs an MCP tool.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;connectMCPTransport&lt;/code&gt; talks to the MCP server without a long-lived env bearer.&lt;/li&gt;
&lt;li&gt;Server demands OAuth. Transport yields a consent challenge.&lt;/li&gt;
&lt;li&gt;Route catches &lt;code&gt;getConsentChallenge&lt;/code&gt;, redirects the browser.&lt;/li&gt;
&lt;li&gt;User consents. Connect issues a runtime token at your callback path.&lt;/li&gt;
&lt;li&gt;Agent retries the tool call. MCP sees a valid OAuth token. Env never held it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If step 5 is “stuff the token into &lt;code&gt;$CONNECT_MCP_TOKEN&lt;/code&gt; so we skip redirects in staging,” staging will ship to production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hard split from the Claude MCP connector
&lt;/h2&gt;

&lt;p&gt;We already covered &lt;a href="https://otf-kit.dev/blog/expo-mcp-connector-claude" rel="noopener noreferrer"&gt;Expo MCP connector for Claude&lt;/a&gt;: EAS builds, Claude, a connector aimed at that editor/runtime pair. This Vercel Connect + TanStack path is a different product surface — different package (&lt;code&gt;@vercel/connect/tanstack-ai&lt;/code&gt;), host (Vercel route + TanStack), auth UX (&lt;code&gt;getConsentChallenge&lt;/code&gt; on &lt;em&gt;your&lt;/em&gt; web route), and secret home (runtime token at the Vercel boundary). Copy-pasting those snippets into a TanStack Vercel app gives you two half-wired OAuth stacks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure modes that are on you
&lt;/h2&gt;

&lt;p&gt;Redirect loops mean the callback did not persist enough session state — fix the session, not the transport. Ignoring &lt;code&gt;getConsentChallenge&lt;/code&gt; so the model “tries another tool” leaks partial work; abort the turn. Log that a challenge happened, not URL query params that may carry one-time codes. Use &lt;code&gt;connectMCPTransport&lt;/code&gt; in local and preview too, or those environments invent env tokens again. Keep AI Gateway credentials (model inference) separate from MCP OAuth (tools) — see &lt;a href="https://otf-kit.dev/blog/mimo-v2-6-ai-gateway-builders" rel="noopener noreferrer"&gt;AI Gateway model picks&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  A minimal implementation checklist
&lt;/h2&gt;

&lt;p&gt;Grounded in the changelog — not invented API:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;toServerSentEventsResponse&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@tanstack/ai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;createMCPClient&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@tanstack/ai-mcp&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;vercelGatewayText&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@tanstack/ai-vercel-gateway&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;connectMCPTransport&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;getConsentChallenge&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@vercel/connect/tanstack-ai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;POST&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;userId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;getUserId&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;linear&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;createMCPClient&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;transport&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;connectMCPTransport&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;http&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;CONNECT_MCP_URL&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;oauth/linear&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;subject&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;userId&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;stream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;adapter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;vercelGatewayText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;anthropic/claude-opus-5&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
      &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;mcp&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;clients&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;linear&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;toServerSentEventsResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;challenge&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;getConsentChallenge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;challenge&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;redirect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;challenge&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;303&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Checklist: &lt;code&gt;@vercel/connect/tanstack-ai&lt;/code&gt; adapter; &lt;code&gt;connectMCPTransport&lt;/code&gt;; catch &lt;code&gt;getConsentChallenge&lt;/code&gt; on the agent route; owned-path consent callback; no MCP access token in env; deny disables tools (no static fallback); logs omit tokens. If the changelog is silent on a step, do not invent it — read &lt;a href="https://vercel.com/changelog/vercel-connect-tanstack-ai" rel="noopener noreferrer"&gt;the entry&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Owned-route auth is the pattern we want when an agent can spend money or read a private repo: the browser user is in the loop, the serverless function is the boundary, and the model is a guest. Install &lt;code&gt;@vercel/connect/tanstack-ai&lt;/code&gt;, send tools through &lt;code&gt;connectMCPTransport&lt;/code&gt;, catch &lt;code&gt;getConsentChallenge&lt;/code&gt; and redirect, let tokens issue at runtime. Leave env for &lt;code&gt;$CONNECT_*&lt;/code&gt; config. Keep the Claude connector in its own post. For a durable product surface under the agent churn, start from the &lt;a href="https://otf-kit.dev/templates" rel="noopener noreferrer"&gt;OTF templates&lt;/a&gt; you own — Connect still owns the OAuth hop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://vercel.com/changelog/vercel-connect-tanstack-ai" rel="noopener noreferrer"&gt;Vercel Connect + TanStack AI changelog&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/gpt-live-1-delegation-ai-gateway" rel="noopener noreferrer"&gt;AI Gateway delegation (OtF)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://otf-kit.dev/blog/expo-mcp-connector-claude" rel="noopener noreferrer"&gt;Expo MCP connector for Claude (OtF)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>vercel</category>
      <category>agents</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Queue backpressure on an owned backend: refuse the burst before the pool melts</title>
      <dc:creator>Dave Kurian</dc:creator>
      <pubDate>Thu, 24 Sep 2026 07:54:05 +0000</pubDate>
      <link>https://dev.to/davekurian/queue-backpressure-on-an-owned-backend-refuse-the-burst-before-the-pool-melts-5h1g</link>
      <guid>https://dev.to/davekurian/queue-backpressure-on-an-owned-backend-refuse-the-burst-before-the-pool-melts-5h1g</guid>
      <description>&lt;p&gt;Agent runs, vendor webhook fan-out, and retry storms arrive as a sudden pile of HTTP calls on an API you operate. The database pool and the worker processes behind that API are a fixed budget. Once every extra call is allowed to pin a connection, the burst becomes latency for every tenant, including the ones that sent a normal amount of traffic.&lt;/p&gt;

&lt;p&gt;Backpressure is the refusal at that edge. Cap in-flight work. Cap how many requests may wait, and for how long. When both caps are exhausted, answer with a status code and a &lt;code&gt;Retry-After&lt;/code&gt; value the caller can honor. A process that keeps an ever-growing list of HTTP requests in memory is betting the pool will catch up. Under this load, it will not.&lt;/p&gt;

&lt;h2&gt;
  
  
  What owned capacity means here
&lt;/h2&gt;

&lt;p&gt;Owned capacity is the part you restart and pay for: database connections, worker processes, and the concurrency budget of the next service you run. The caller does not know those numbers. An agent loop will keep issuing work. Admission turns your budget into a yes, a short wait, or a reject on the backend that holds the pool.&lt;/p&gt;

&lt;p&gt;Matthew Palma's notes on HTTP API admission control describe that gate as a concurrency cap, a bounded queue with a maximum wait, and a reject that carries 503 and &lt;code&gt;Retry-After&lt;/code&gt;. Readiness fails when the process should leave rotation. Size the knobs from your pool: &lt;code&gt;$DB_POOL_SIZE&lt;/code&gt;, &lt;code&gt;$MAX_INFLIGHT&lt;/code&gt;, &lt;code&gt;$MAX_QUEUE_DEPTH&lt;/code&gt;, &lt;code&gt;$QUEUE_WAIT_MS&lt;/code&gt;, and &lt;code&gt;$RETRY_AFTER_SEC&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;API gateway guidance from API7 separates two knobs dashboards often mash together. A rate limit counts arrivals in a window. Concurrency counts work still inside the system. A client under its rate limit can still occupy every database connection when each call is slow. &lt;a href="https://otf-kit.dev/blog/rate-limiting-owned-api-production" rel="noopener noreferrer"&gt;Rate limiting an owned API&lt;/a&gt; is the arrival window. This piece is occupancy. You want the pair.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the other queues in their own posts
&lt;/h2&gt;

&lt;p&gt;This series already uses "queue" for three different jobs. Mixing them buffers the wrong thing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://otf-kit.dev/blog/circuit-breakers-owned-backend" rel="noopener noreferrer"&gt;Circuit breakers on an owned backend&lt;/a&gt; fail fast on outbound calls when a vendor is already failing. Inbound admission decides whether a new request may enter your database at all.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://otf-kit.dev/blog/dead-letter-queue-background-jobs" rel="noopener noreferrer"&gt;Dead letter queues for background jobs&lt;/a&gt; quarantine a message after its retries are spent. That storage sits later than the HTTP gate. A request you have not admitted yet has nothing to dead-letter.&lt;/p&gt;

&lt;p&gt;A client-side offline queue holds mutations on the device and replays them when the network returns. The replay still has to pass the server gate, or the reconnect wave becomes the burst this post is about.&lt;/p&gt;

&lt;p&gt;An unbounded in-process queue stores the overload until memory, file descriptors, or the database pool fail together. Bounded concurrency plus a short wait, then 429 or 503 with &lt;code&gt;Retry-After&lt;/code&gt;, spends the pool on work already admitted and tells everyone else to come back.&lt;/p&gt;

&lt;h2&gt;
  
  
  Admit, wait briefly, or reject
&lt;/h2&gt;

&lt;p&gt;Per process, or per worker behind a load balancer, take a semaphore of size &lt;code&gt;$MAX_INFLIGHT&lt;/code&gt;. Derive it from &lt;code&gt;$DB_POOL_SIZE&lt;/code&gt; and any downstream budget the handler also consumes. If each admitted request holds one pool connection for the life of the call, &lt;code&gt;$MAX_INFLIGHT&lt;/code&gt; must stay inside what that pool can finish. Extra handlers wait inside the driver, and the reject you wanted becomes a timeout.&lt;/p&gt;

&lt;p&gt;When several workers share one database, the fleet total is the real cap. Four processes each sized at the full &lt;code&gt;$DB_POOL_SIZE&lt;/code&gt; admit four times what the pool can serve. Divide the pool across the processes, and leave a remainder for migrations, health checks, and admin work.&lt;/p&gt;

&lt;p&gt;In front of the semaphore, a wait queue is optional and short. Cap it with &lt;code&gt;$MAX_QUEUE_DEPTH&lt;/code&gt; and &lt;code&gt;$QUEUE_WAIT_MS&lt;/code&gt;. When the queue is full, or a waiter exceeds the wait budget, reject without checking out a connection.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxke3uhbo59sdu2spanqt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxke3uhbo59sdu2spanqt.png" alt="Bounded concurrency dial and short wait queue with Retry-After tokens" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;API7 frames that choice as bounded delay or reject, and names load shedding for the moment shared capacity is already gone.&lt;/p&gt;

&lt;p&gt;429 is the caller or the tenant. This API key, this agent, this customer has too much in flight relative to the share you assigned them. &lt;a href="https://www.rfc-editor.org/rfc/rfc6585#section-4" rel="noopener noreferrer"&gt;RFC 6585 section 4&lt;/a&gt; defines 429 Too Many Requests and allows a &lt;code&gt;Retry-After&lt;/code&gt; header so the client knows when another attempt is reasonable.&lt;/p&gt;

&lt;p&gt;503 is shared capacity. The process, the pool, or the worker fleet is full no matter who called. &lt;a href="https://www.rfc-editor.org/rfc/rfc9110#name-retry-after" rel="noopener noreferrer"&gt;RFC 9110&lt;/a&gt; specifies &lt;code&gt;Retry-After&lt;/code&gt; as either a delay in seconds or an HTTP-date. For this gate, send &lt;code&gt;Retry-After: $RETRY_AFTER_SEC&lt;/code&gt; on both 429 and 503. Callers that honor the header back off. Callers that ignore it still hit the same cheap reject on the next try.&lt;/p&gt;

&lt;p&gt;Do not answer the shed with a 200 whose body claims the work was queued while the process has nowhere to run it. That hides the overload and invites the client to poll.&lt;/p&gt;

&lt;h2&gt;
  
  
  Occupancy moves when latency moves
&lt;/h2&gt;

&lt;p&gt;API7 points at Little's Law for a first estimate: average concurrency is about arrival rate multiplied by time in the system. When a dependency slows and handler latency doubles, the same arrival rate doubles how many calls are in flight. The rate limit still says the tenant is fine. The semaphore sees the extra occupancy and starts to wait or reject.&lt;/p&gt;

&lt;p&gt;That is why &lt;code&gt;$MAX_INFLIGHT&lt;/code&gt; is tied to &lt;code&gt;$DB_POOL_SIZE&lt;/code&gt;, then revised when handler time changes. Keep the rate limit beside the semaphore: rate stops one tenant from filling &lt;code&gt;$MAX_QUEUE_DEPTH&lt;/code&gt; with cheap arrivals; concurrency stops a slow dependency from holding &lt;code&gt;$DB_POOL_SIZE&lt;/code&gt; until every other request times out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Metrics, readiness, and the path
&lt;/h2&gt;

&lt;p&gt;Emit &lt;code&gt;queue_depth&lt;/code&gt; and &lt;code&gt;admission_rejected&lt;/code&gt;. Split rejects by 429 and 503 if both exist. Shedding is the gate working. A dashboard that only alerts on 500 stays green while you turn traffic away.&lt;/p&gt;

&lt;p&gt;Fail readiness when &lt;code&gt;queue_depth&lt;/code&gt; or processing lag crosses the threshold for this process. Readiness is what the platform checks before it sends new work. Liveness can remain true so the process finishes requests it already admitted. Palma includes readiness in the admission story for that reason: a live, saturated process should drop out of rotation until the wait queue recedes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://otf-kit.dev/blog/graceful-shutdown-drain-owned-backend" rel="noopener noreferrer"&gt;Draining an owned backend on shutdown&lt;/a&gt; is the same refusal with a different trigger. Shutdown stops new work, then waits for in-flight handlers to finish. Backpressure stops new work while the process is supposed to stay up, because the pool is already committed.&lt;/p&gt;

&lt;p&gt;The path is a straight line. In-flight below &lt;code&gt;$MAX_INFLIGHT&lt;/code&gt; means the handler runs. In-flight full, and queue depth below &lt;code&gt;$MAX_QUEUE_DEPTH&lt;/code&gt;, means the request waits up to &lt;code&gt;$QUEUE_WAIT_MS&lt;/code&gt;. Wait expired, or queue full, means 429 or 503 with &lt;code&gt;Retry-After&lt;/code&gt;, and the pool is never checked out.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3cxuhxwanarzt6uz4l2b.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3cxuhxwanarzt6uz4l2b.png" alt="Admit, bounded wait, and reject path keeping the database pool healthy" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A gate in front of the handler
&lt;/h2&gt;

&lt;p&gt;One process-wide semaphore, one bounded waiter list, one reject response. A per-tenant cap uses the same shape under the process cap. Scope &lt;code&gt;"tenant"&lt;/code&gt; maps to 429. Scope &lt;code&gt;"process"&lt;/code&gt; maps to 503. Call &lt;code&gt;release&lt;/code&gt; in a &lt;code&gt;finally&lt;/code&gt;. Skip the database until &lt;code&gt;admit&lt;/code&gt; returns &lt;code&gt;ok&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Pseudocode: semaphore + bounded waiters + reject&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;admit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;tenant&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;process&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Admit&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;inflight&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nc"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;MAX_INFLIGHT&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;inflight&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;release&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;queueDepth&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="nc"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;MAX_QUEUE_DEPTH&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;scope&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;tenant&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="mi"&gt;429&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;503&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;retryAfterSec&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;RETRY_AFTER_SEC&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="c1"&gt;// wait up to QUEUE_WAIT_MS, then same reject shape&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Increment &lt;code&gt;admission_rejected&lt;/code&gt; whenever &lt;code&gt;admit&lt;/code&gt; returns &lt;code&gt;ok: false&lt;/code&gt;, and return that status with &lt;code&gt;Retry-After: $RETRY_AFTER_SEC&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Broker backlog vs HTTP buffer
&lt;/h2&gt;

&lt;p&gt;A durable job broker with retention and a consumer concurrency setting is a real queue. Consumer count still has to respect &lt;code&gt;$DB_POOL_SIZE&lt;/code&gt;, or the jobs pin the same pool the HTTP gate just protected. An array on the API process has no retention and no bound until you add &lt;code&gt;$MAX_QUEUE_DEPTH&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Agent clients retry on their own timer. Every extra attempt still gets a cheap reject: no pool checkout. Log &lt;code&gt;admission_rejected&lt;/code&gt; and the scope.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checklist
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Set &lt;code&gt;$MAX_INFLIGHT&lt;/code&gt; per process from &lt;code&gt;$DB_POOL_SIZE&lt;/code&gt;. Across workers, the sum still has to fit the pool.&lt;/li&gt;
&lt;li&gt;Cap any wait queue with &lt;code&gt;$MAX_QUEUE_DEPTH&lt;/code&gt; and &lt;code&gt;$QUEUE_WAIT_MS&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;On a full gate, return 429 for a caller or tenant budget and 503 for shared capacity, both with &lt;code&gt;Retry-After: $RETRY_AFTER_SEC&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Emit &lt;code&gt;queue_depth&lt;/code&gt; and &lt;code&gt;admission_rejected&lt;/code&gt;. Fail readiness when depth or lag crosses the threshold.&lt;/li&gt;
&lt;li&gt;Keep a rate limit on arrivals. Keep HTTP request buffering bounded.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scaffolding the API, the worker, and the repo you deploy is a separate decision from this gate. OTF ships full-stack kits you own — Booking, Fitness, and SaaS Dashboard at $99 each, or the Everything Bundle at $149 — plus a free SDK under MIT. The templates live at &lt;a href="https://otf-kit.dev/templates" rel="noopener noreferrer"&gt;otf-kit.dev/templates&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Protect owned capacity under load with a hard in-flight cap, a short wait budget, and an honest 429 or 503 with &lt;code&gt;Retry-After&lt;/code&gt; at your edge. Infinite buffering is not backpressure. It is a deferred outage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.rfc-editor.org/rfc/rfc6585#section-4" rel="noopener noreferrer"&gt;RFC 6585 §4, 429 Too Many Requests&lt;/a&gt; — defines 429 and allows Retry-After on that response.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.rfc-editor.org/rfc/rfc9110#name-retry-after" rel="noopener noreferrer"&gt;RFC 9110, Retry-After&lt;/a&gt; — delay in seconds or an HTTP-date; the time to wait before the next request.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://matthewpalma.dev/blog/http-api-admission-control-concurrency-queues-load-shedding" rel="noopener noreferrer"&gt;Matthew Palma, HTTP API admission control&lt;/a&gt; — concurrency caps, a bounded queue with maxWaitMs, 503 plus Retry-After, and readiness.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://api7.ai/learning-center/api-gateway-guide/api-gateway-concurrency-control" rel="noopener noreferrer"&gt;API7, API gateway concurrency control&lt;/a&gt; — rate versus concurrency, Little's Law as an estimate, 429 for a caller or tenant versus 503 for shared capacity, bounded delay or reject, and load shedding.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>architecture</category>
      <category>backend</category>
      <category>agents</category>
    </item>
    <item>
      <title>Circuit breakers on an owned backend: fail fast when the AI vendor fails</title>
      <dc:creator>Dave Kurian</dc:creator>
      <pubDate>Thu, 24 Sep 2026 05:31:34 +0000</pubDate>
      <link>https://dev.to/davekurian/circuit-breakers-on-an-owned-backend-fail-fast-when-the-ai-vendor-fails-4bnj</link>
      <guid>https://dev.to/davekurian/circuit-breakers-on-an-owned-backend-fail-fast-when-the-ai-vendor-fails-4bnj</guid>
      <description>&lt;p&gt;When a downstream AI or vendor HTTP dependency starts failing, the owned API should stop waiting on it. A circuit breaker opens after a measured run of failures and returns a cached, stale, or degraded response immediately. The process stays available. Threads and event-loop slots are not pinned to a vendor that is already down. Retries alone do the opposite: every caller keeps burning the timeout budget and the retry budget, which multiplies load on a dependency that is already sick.&lt;/p&gt;

&lt;p&gt;This post is the breaker state machine. Outbound wait and retry live in &lt;a href="https://otf-kit.dev/blog/api-timeouts-retries-ai-backends" rel="noopener noreferrer"&gt;timeouts and retries for AI backends&lt;/a&gt;. Parking work for later is a &lt;a href="https://otf-kit.dev/blog/dead-letter-queue-background-jobs" rel="noopener noreferrer"&gt;dead-letter queue&lt;/a&gt; concern. The breaker decides whether this dependency may be called right now.&lt;/p&gt;

&lt;h2&gt;
  
  
  What cascading failure looks like without a breaker
&lt;/h2&gt;

&lt;p&gt;Martin Fowler's circuit breaker note and Microsoft's Azure Architecture Center pattern describe the same operator failure mode. A remote call hangs or errors. Callers retry. Each attempt occupies a worker or a promise until the client timeout fires. The owned API's latency SLO collapses even though its own code is fine, because capacity is stuck on calls that cannot succeed.&lt;/p&gt;

&lt;p&gt;AI providers make this sharp: a chat or embeddings call can sit for tens of seconds before a 503, a 429, or a transport error. Retries multiply in-flight work. Health checks that also call the vendor start failing. Autoscaling adds instances that open more connections to the same broken endpoint. A breaker does not fix the vendor. It bounds the blast radius so routes you still own — auth, database reads, previously computed answers — keep returning.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp9shjlzvvluzvdvqniri.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp9shjlzvvluzvdvqniri.png" alt="Closed, open, and half-open circuit breaker paths with fallback cache" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Closed, open, half-open
&lt;/h2&gt;

&lt;p&gt;The state machine used by Azure Architecture Center, AWS Prescriptive Guidance, Resilience4j, and the Node and Go libraries below is the same three states Fowler sketched.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Closed.&lt;/strong&gt; Calls go through. Outcomes are recorded. Failures count only when they are dependency failures. If the failure rate over the window stays under the threshold, the breaker stays closed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Open.&lt;/strong&gt; Calls are not attempted. The owned handler returns the fallback immediately and logs that the call was rejected because the circuit is open. A timer (&lt;code&gt;$CB_WAIT_OPEN_MS&lt;/code&gt;) must elapse before any probe is allowed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Half-open.&lt;/strong&gt; After the wait, a small number of trial calls (&lt;code&gt;$CB_HALF_OPEN_CALLS&lt;/code&gt;) are permitted. Success closes the breaker. Failure opens it again and restarts the wait. Half-open discovers recovery without dumping full QPS onto a vendor that just came back.&lt;/p&gt;

&lt;p&gt;While open, the fallback is a product decision: cached completion, last-known embeddings, a static degraded payload, or a 503 with a short &lt;code&gt;Retry-After&lt;/code&gt; you control. You do not learn the vendor is still down by waiting for &lt;code&gt;$AI_PROVIDER_BASE_URL&lt;/code&gt; to time out again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why retries without a breaker amplify load
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffq4k9gldq6krqxlv0mjz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffq4k9gldq6krqxlv0mjz.png" alt="Retries-only pileup versus breaker diverting traffic to fallback" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Retries only&lt;/th&gt;
&lt;th&gt;Breaker plus fallback&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Vendor 503 storm&lt;/td&gt;
&lt;td&gt;Every request waits, then retries&lt;/td&gt;
&lt;td&gt;After the threshold, new requests skip the network&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;In-flight work&lt;/td&gt;
&lt;td&gt;Grows with concurrency × attempts&lt;/td&gt;
&lt;td&gt;Caps at calls already in flight when the breaker opens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User latency&lt;/td&gt;
&lt;td&gt;Sum of timeouts and backoffs&lt;/td&gt;
&lt;td&gt;Fallback latency, typically a cache read&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recovery&lt;/td&gt;
&lt;td&gt;All callers slam the vendor on the first green blip&lt;/td&gt;
&lt;td&gt;Half-open probes, then close&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4xx from your bug&lt;/td&gt;
&lt;td&gt;Often retried if classified poorly&lt;/td&gt;
&lt;td&gt;Ignored by the breaker; you still fix the client&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Azure's pattern page is explicit: a circuit breaker is not a substitute for retry. Retries and timeouts wrap one attempt sequence; if the breaker is open, do not start that sequence. &lt;code&gt;CallNotPermittedException&lt;/code&gt; (Resilience4j) and opossum's &lt;code&gt;reject&lt;/code&gt; event are the signals to stop.&lt;/p&gt;

&lt;p&gt;A DLQ is the wrong tool for this request path — it parks work you already accepted. A user on &lt;code&gt;POST /v1/answer&lt;/code&gt; needs fail-fast plus degraded body; background enrichment can use the DLQ separately.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to count as failure
&lt;/h2&gt;

&lt;p&gt;Resilience4j documents a 50 percent failure-rate threshold, sliding windows, &lt;code&gt;minimumNumberOfCalls&lt;/code&gt;, slow-call rate, wait in open state, half-open permits, and &lt;code&gt;recordExceptions&lt;/code&gt; / &lt;code&gt;ignoreExceptions&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Count toward the breaker: outbound client timeouts; connection resets, DNS failures, TLS handshake errors; HTTP 5xx; HTTP 429 and 503 treated as dependency health signals.&lt;/p&gt;

&lt;p&gt;Do not count: HTTP 4xx that are your bug or bad user input (400, 404, 422); client cancellations; business outcomes you already decided. Opening on 4xx lets a malformed-body deploy take the dependency "down" while the vendor is healthy.&lt;/p&gt;

&lt;p&gt;Slow calls deserve a separate rate: a body that always arrives in 28 seconds will not trip failure-rate and will still exhaust your pool. Use Resilience4j's slow-call threshold or opossum's &lt;code&gt;timeout&lt;/code&gt; — pick one definition of "too slow."&lt;/p&gt;

&lt;h2&gt;
  
  
  One breaker per dependency
&lt;/h2&gt;

&lt;p&gt;Do not share one breaker across unrelated providers or shards. A search vendor and a chat vendor fail independently. Key by dependency name: &lt;code&gt;chat-provider&lt;/code&gt;, &lt;code&gt;embeddings-provider&lt;/code&gt;, not &lt;code&gt;outbound-http&lt;/code&gt;. Extra breakers are cheap; a false open on unrelated traffic is not.&lt;/p&gt;

&lt;p&gt;Name them in metrics and logs. On transition and reject, log breaker name, new state, correlation id, and route — the same discipline as &lt;a href="https://otf-kit.dev/blog/production-structured-logging-for-agents" rel="noopener noreferrer"&gt;production structured logging for agents&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Libraries and settings
&lt;/h2&gt;

&lt;p&gt;Prefer a maintained implementation. &lt;strong&gt;Resilience4j&lt;/strong&gt; (JVM): sliding windows, &lt;code&gt;failureRateThreshold&lt;/code&gt; default 50, slow-call rate, minimum calls, open wait, half-open permits, &lt;code&gt;CallNotPermittedException&lt;/code&gt;, exception filters. &lt;strong&gt;opossum&lt;/strong&gt; (Node): &lt;code&gt;timeout&lt;/code&gt;, &lt;code&gt;errorThresholdPercentage&lt;/code&gt;, &lt;code&gt;resetTimeout&lt;/code&gt;, &lt;code&gt;fallback&lt;/code&gt;, plus &lt;code&gt;reject&lt;/code&gt;/&lt;code&gt;open&lt;/code&gt;/&lt;code&gt;halfOpen&lt;/code&gt;/&lt;code&gt;close&lt;/code&gt; events. &lt;strong&gt;sony/gobreaker&lt;/strong&gt; (Go): ready-to-trip, open-state timeout, max half-open requests. AWS Prescriptive Guidance and Microsoft Azure Architecture Center describe the same machine without a library.&lt;/p&gt;

&lt;p&gt;Read at process start: &lt;code&gt;$CB_FAILURE_RATE_THRESHOLD&lt;/code&gt;, &lt;code&gt;$CB_SLIDING_WINDOW&lt;/code&gt;, &lt;code&gt;$CB_WAIT_OPEN_MS&lt;/code&gt;, &lt;code&gt;$CB_HALF_OPEN_CALLS&lt;/code&gt;, and &lt;code&gt;$AI_PROVIDER_BASE_URL&lt;/code&gt;. Require a minimum number of calls before evaluating rate — opening on the first failure flaps.&lt;/p&gt;

&lt;h2&gt;
  
  
  TypeScript sketch for one outbound AI client
&lt;/h2&gt;

&lt;p&gt;The timeout and retry wrapper is the one from the timeouts post. The breaker sits outside it so an open circuit never enters the retry loop.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;AiResult&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;unknown&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;servedBy&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;live&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;cache&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;degraded&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;circuit_open&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;upstream_failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="c1"&gt;// settings: $CB_FAILURE_RATE_THRESHOLD $CB_SLIDING_WINDOW&lt;/span&gt;
&lt;span class="c1"&gt;// $CB_WAIT_OPEN_MS $CB_HALF_OPEN_CALLS ; URL: $AI_PROVIDER_BASE_URL&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;completeWithBreaker&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;correlationId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;promptHash&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;AiResult&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;breaker&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;breakers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;chat-provider&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;tryAcquire&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;warn&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;circuit_rejected&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;chat-provider&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;state&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;state&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
      &lt;span class="nx"&gt;correlationId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cached&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;promptHash&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cached&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;cached&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;servedBy&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;cache&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;degraded&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;circuit_open&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;withTimeoutAndRetry&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt;
      &lt;span class="nf"&gt;postChat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;$AI_PROVIDER_BASE_URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nx"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recordSuccess&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;promptHash&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;servedBy&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;live&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;isDependencyFailure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="nx"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recordFailure&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="nx"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recordSuccess&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// 4xx business: not vendor health&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cached&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;promptHash&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cached&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;cached&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;servedBy&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;cache&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;degraded&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;upstream_failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;tryAcquire&lt;/code&gt; returns false when open or when half-open already has &lt;code&gt;$CB_HALF_OPEN_CALLS&lt;/code&gt; in flight. Count timeouts, connection errors, 429, 503, and other 5xx as failures; ignore 4xx for the window. Invoke &lt;code&gt;withTimeoutAndRetry&lt;/code&gt; only after acquire succeeds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checklist for the owned API
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;One breaker per independent dependency, named in logs and metrics.&lt;/li&gt;
&lt;li&gt;Record failures for timeouts, connection errors, 5xx, and vendor 429/503.&lt;/li&gt;
&lt;li&gt;Ignore 4xx that indicate a bad request rather than a bad dependency.&lt;/li&gt;
&lt;li&gt;Evaluate failure rate only after a minimum number of calls in &lt;code&gt;$CB_SLIDING_WINDOW&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Open when the rate exceeds &lt;code&gt;$CB_FAILURE_RATE_THRESHOLD&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;While open, skip &lt;code&gt;$AI_PROVIDER_BASE_URL&lt;/code&gt;; return cache or degraded body; log breaker name, state, correlation id.&lt;/li&gt;
&lt;li&gt;After &lt;code&gt;$CB_WAIT_OPEN_MS&lt;/code&gt;, allow &lt;code&gt;$CB_HALF_OPEN_CALLS&lt;/code&gt; probes; close on success, re-open on failure.&lt;/li&gt;
&lt;li&gt;Emit metrics on every transition and reject.&lt;/li&gt;
&lt;li&gt;Keep timeouts and a small retry budget for the closed state. Stop retrying when the call is not permitted.&lt;/li&gt;
&lt;li&gt;Do not use the breaker as a DLQ, and do not use a DLQ as a live fail-fast.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  How this sits next to the other controls
&lt;/h2&gt;

&lt;p&gt;Timeouts cap one call. Retries spend a fixed budget on transient faults. The breaker removes the dependency from the hot path after that budget is clearly wasted. Rate limiting protects you from your callers; it does not protect you from a vendor that accepts the connection and then stalls. Background jobs that call the same vendor should use a sibling breaker with a looser threshold so batch volume does not open the circuit for interactive traffic — when retries are spent and the breaker is open, park the job in the DLQ.&lt;/p&gt;

&lt;p&gt;While open, return a cache hit marked &lt;code&gt;servedBy: "cache"&lt;/code&gt;, a stale-but-bounded row, a stable degraded JSON object with your correlation id, or a null optional enrichment while the primary database read still returns. Do not fall through an open breaker into an unbounded call to a second vendor — give the backup its own breaker.&lt;/p&gt;

&lt;p&gt;If you own the repo, put the breaker next to the HTTP client that already holds timeouts and structured log fields. OTF's owned-repo kits are built around that client boundary; add the breaker in the same module instead of inventing a second outbound stack.&lt;/p&gt;

&lt;p&gt;The vendor will fail again. The owned API should notice quickly, stop calling, serve the fallback, and try a few probes later. That is the whole machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Martin Fowler, CircuitBreaker: &lt;a href="https://martinfowler.com/bliki/CircuitBreaker.html" rel="noopener noreferrer"&gt;https://martinfowler.com/bliki/CircuitBreaker.html&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Azure Architecture Center, Circuit Breaker pattern: &lt;a href="https://learn.microsoft.com/en-us/azure/architecture/patterns/circuit-breaker" rel="noopener noreferrer"&gt;https://learn.microsoft.com/en-us/azure/architecture/patterns/circuit-breaker&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Resilience4j CircuitBreaker: &lt;a href="https://resilience4j.readme.io/docs/circuitbreaker" rel="noopener noreferrer"&gt;https://resilience4j.readme.io/docs/circuitbreaker&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;AWS Prescriptive Guidance, Circuit breaker pattern: &lt;a href="https://docs.aws.amazon.com/prescriptive-guidance/latest/cloud-design-patterns/circuit-breaker.html" rel="noopener noreferrer"&gt;https://docs.aws.amazon.com/prescriptive-guidance/latest/cloud-design-patterns/circuit-breaker.html&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;opossum (Node): &lt;a href="https://www.npmjs.com/package/opossum" rel="noopener noreferrer"&gt;https://www.npmjs.com/package/opossum&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;sony/gobreaker (Go): &lt;a href="https://github.com/sony/gobreaker" rel="noopener noreferrer"&gt;https://github.com/sony/gobreaker&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Related: &lt;a href="https://otf-kit.dev/blog/api-timeouts-retries-ai-backends" rel="noopener noreferrer"&gt;https://otf-kit.dev/blog/api-timeouts-retries-ai-backends&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Related: &lt;a href="https://otf-kit.dev/blog/dead-letter-queue-background-jobs" rel="noopener noreferrer"&gt;https://otf-kit.dev/blog/dead-letter-queue-background-jobs&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Related: &lt;a href="https://otf-kit.dev/blog/production-structured-logging-for-agents" rel="noopener noreferrer"&gt;https://otf-kit.dev/blog/production-structured-logging-for-agents&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>architecture</category>
      <category>backend</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
