<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ray Mac</title>
    <description>The latest articles on DEV Community by Ray Mac (@ray_mac).</description>
    <link>https://dev.to/ray_mac</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4107278%2Fb99cc365-5c7f-428d-975a-436dff513899.jpg</url>
      <title>DEV Community: Ray Mac</title>
      <link>https://dev.to/ray_mac</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ray_mac"/>
    <language>en</language>
    <item>
      <title>Transcription Pricing Compared (2026): ScribeToAny vs Otter, Rev, Sonix and HappyScribe</title>
      <dc:creator>Ray Mac</dc:creator>
      <pubDate>Sun, 27 Sep 2026 14:42:27 +0000</pubDate>
      <link>https://dev.to/ray_mac/transcription-pricing-compared-2026-scribetoany-vs-otter-rev-sonix-and-happyscribe-1bb5</link>
      <guid>https://dev.to/ray_mac/transcription-pricing-compared-2026-scribetoany-vs-otter-rev-sonix-and-happyscribe-1bb5</guid>
      <description>&lt;p&gt;Transcription tools love to advertise a monthly price, but the monthly price&lt;br&gt;
tells you very little. What matters is &lt;strong&gt;how many hours of audio that price&lt;br&gt;
covers&lt;/strong&gt;, and whether the plan lets you upload the files you actually have.&lt;br&gt;
Two tools at the same $20 can differ tenfold in what you get.&lt;/p&gt;

&lt;p&gt;So instead of comparing sticker prices, this post works out what a real&lt;br&gt;
workload costs on five tools: &lt;strong&gt;Otter.ai, Rev, Sonix, HappyScribe and&lt;br&gt;
ScribeToAny&lt;/strong&gt; (ours — we'll be upfront about where the others are the better&lt;br&gt;
choice). Every number comes from each vendor's public pricing page, checked on&lt;br&gt;
&lt;strong&gt;September 23, 2026&lt;/strong&gt;; links are at the end. Prices change, so treat this as a&lt;br&gt;
snapshot and check the vendor's page before you buy.&lt;/p&gt;

&lt;p&gt;One scoping note: this is about &lt;strong&gt;transcribing recorded files&lt;/strong&gt; — interviews,&lt;br&gt;
lectures, podcasts, videos. Live meeting bots are a different job, covered&lt;br&gt;
under "When another tool is the better choice" below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost per hour of audio
&lt;/h2&gt;

&lt;p&gt;Each row is the cheapest plan that includes file uploads, divided by the hours&lt;br&gt;
of transcription it includes each month.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool and plan&lt;/th&gt;
&lt;th&gt;Price&lt;/th&gt;
&lt;th&gt;Included per month&lt;/th&gt;
&lt;th&gt;Cost per hour&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ScribeToAny Plus&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$10/mo · $88/yr&lt;/td&gt;
&lt;td&gt;25 h (1,500 min)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;$0.40&lt;/strong&gt; · $0.29 yearly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ScribeToAny Pro&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$20/mo · $120/yr&lt;/td&gt;
&lt;td&gt;50 h (3,000 min)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;$0.40&lt;/strong&gt; · $0.20 yearly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rev Essentials&lt;/td&gt;
&lt;td&gt;$29.99/mo · $305.90/yr&lt;/td&gt;
&lt;td&gt;83 h (5,000 min)&lt;/td&gt;
&lt;td&gt;$0.36 · $0.31 yearly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Otter.ai Pro&lt;/td&gt;
&lt;td&gt;$16.99/mo · $99.96/yr&lt;/td&gt;
&lt;td&gt;20 h (1,200 min), &lt;strong&gt;max 10 file imports&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;$0.85 · $0.42 yearly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HappyScribe Pro&lt;/td&gt;
&lt;td&gt;€29/mo · €228/yr&lt;/td&gt;
&lt;td&gt;10 h (600 min)&lt;/td&gt;
&lt;td&gt;€2.90 · €1.90 yearly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HappyScribe Basic&lt;/td&gt;
&lt;td&gt;€17/mo · €102/yr&lt;/td&gt;
&lt;td&gt;2 h (120 min)&lt;/td&gt;
&lt;td&gt;€8.50 · €4.25 yearly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonix Core&lt;/td&gt;
&lt;td&gt;$25/mo · $275/yr&lt;/td&gt;
&lt;td&gt;5 h&lt;/td&gt;
&lt;td&gt;$5.00 · $4.58 yearly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonix pay-as-you-go&lt;/td&gt;
&lt;td&gt;$10 per hour&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things stand out. Per hour, &lt;strong&gt;Rev Essentials is slightly cheaper than&lt;br&gt;
ScribeToAny on monthly billing&lt;/strong&gt; ($0.36 vs $0.40; on yearly billing ScribeToAny&lt;br&gt;
is cheaper) — but it starts at $29.99 and covers English and Spanish only;&lt;br&gt;
you need Rev Pro ($59.99/mo) for its 37+ languages. And &lt;strong&gt;Otter's 1,200&lt;br&gt;
minutes are capped at 10 imported files a month&lt;/strong&gt; on Pro, so if your work is&lt;br&gt;
uploading recordings rather than recording meetings, the file cap runs out long&lt;br&gt;
before the minutes do.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a real month costs
&lt;/h2&gt;

&lt;p&gt;Cost per hour hides the jumps between plans, so here are two concrete&lt;br&gt;
workloads, priced on monthly billing (yearly in brackets).&lt;/p&gt;

&lt;h3&gt;
  
  
  10 hours of recordings a month
&lt;/h3&gt;

&lt;p&gt;Ten one-hour interviews, or twenty 30-minute lectures.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Plan you'd need&lt;/th&gt;
&lt;th&gt;Monthly cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ScribeToAny&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Plus (1,500 min)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;$10&lt;/strong&gt; ($7.33)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Otter.ai&lt;/td&gt;
&lt;td&gt;Pro — fits only if it's ≤ 10 files&lt;/td&gt;
&lt;td&gt;$16.99 ($8.33)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rev&lt;/td&gt;
&lt;td&gt;Essentials&lt;/td&gt;
&lt;td&gt;$29.99 ($25.49)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HappyScribe&lt;/td&gt;
&lt;td&gt;Pro (600 min)&lt;/td&gt;
&lt;td&gt;€29 (€19)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonix&lt;/td&gt;
&lt;td&gt;Core 5 h + 5 h at $10/h&lt;/td&gt;
&lt;td&gt;$75&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  30 hours of recordings a month
&lt;/h3&gt;

&lt;p&gt;A research project, a weekly podcast back catalogue, or a course's lectures.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Plan you'd need&lt;/th&gt;
&lt;th&gt;Monthly cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ScribeToAny&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Pro (3,000 min)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;$20&lt;/strong&gt; ($10)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rev&lt;/td&gt;
&lt;td&gt;Essentials (5,000 min)&lt;/td&gt;
&lt;td&gt;$29.99 ($25.49)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Otter.ai&lt;/td&gt;
&lt;td&gt;Business — Pro stops at 1,200 min and 10 files&lt;/td&gt;
&lt;td&gt;$30 ($19.99)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HappyScribe&lt;/td&gt;
&lt;td&gt;Business (6,000 min)&lt;/td&gt;
&lt;td&gt;€89 (€59)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonix&lt;/td&gt;
&lt;td&gt;Pro (40 h)&lt;/td&gt;
&lt;td&gt;$80 ($73.33)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At 30 hours, the gap with Sonix and HappyScribe is several-fold. Rev and Otter&lt;br&gt;
Business come closest, at roughly $20–30 a month against ScribeToAny's $10–20 —&lt;br&gt;
and Otter Business is a per-user plan built for meetings rather than file&lt;br&gt;
uploads.&lt;/p&gt;

&lt;h2&gt;
  
  
  Free plans: what you can actually do before paying
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Free allowance&lt;/th&gt;
&lt;th&gt;Longest file&lt;/th&gt;
&lt;th&gt;Exports on free&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ScribeToAny&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;100 min every month, 2 files/day&lt;/td&gt;
&lt;td&gt;30 min&lt;/td&gt;
&lt;td&gt;All 8: TXT, SRT, VTT, TSV, CSV, JSON, PDF, DOCX&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Otter.ai&lt;/td&gt;
&lt;td&gt;300 min/month, but &lt;strong&gt;3 file imports in total&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;30 min&lt;/td&gt;
&lt;td&gt;TXT, MP3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rev&lt;/td&gt;
&lt;td&gt;45 min/month, English only&lt;/td&gt;
&lt;td&gt;not listed&lt;/td&gt;
&lt;td&gt;not listed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HappyScribe&lt;/td&gt;
&lt;td&gt;10-minute one-time trial&lt;/td&gt;
&lt;td&gt;45 min&lt;/td&gt;
&lt;td&gt;TXT, SRT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonix&lt;/td&gt;
&lt;td&gt;30-minute one-time trial&lt;/td&gt;
&lt;td&gt;not listed&lt;/td&gt;
&lt;td&gt;not listed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;ScribeToAny's free plan is smaller than Otter's on paper, but it renews every&lt;br&gt;
month, has no lifetime import cap, and includes speaker labels and every export&lt;br&gt;
format — so you can test your real files, in the format you need, before paying&lt;br&gt;
anything. No credit card is required.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits that change the math
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;File length.&lt;/strong&gt; ScribeToAny paid plans take files up to &lt;strong&gt;3 hours&lt;/strong&gt; (5 GB).
Otter Pro stops at 90 minutes per conversation and HappyScribe Basic at 90
minutes per file; long recordings would need splitting there.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Languages.&lt;/strong&gt; ScribeToAny transcribes about &lt;strong&gt;98 languages&lt;/strong&gt; on every plan,
free included, and can transcribe non-English audio straight into English.
Otter supports 6 languages; Rev Essentials 2 (37+ on Pro); Sonix 54+;
HappyScribe 150+ (60+ on its free trial).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parallel work.&lt;/strong&gt; ScribeToAny paid plans transcribe 6 files at once, with no
daily upload limit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Translation.&lt;/strong&gt; Finished ScribeToAny transcripts can be translated into 134+
languages for free (via Google Translate), keeping the original subtitle
timings.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  When another tool is the better choice
&lt;/h2&gt;

&lt;p&gt;Cheaper isn't better if the tool doesn't do the job. Pick one of the others if:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You need live meeting notes.&lt;/strong&gt; Otter joins Zoom, Teams and Google Meet and
transcribes in real time. ScribeToAny transcribes recordings you upload; it
doesn't join calls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You need human-verified transcripts.&lt;/strong&gt; Rev offers human transcription at
$1.99/min with 99%+ accuracy, and legal transcript formats — useful for court
and compliance work. ScribeToAny is machine transcription only; see
&lt;a href="https://scribetoany.com/blog/how-accurate-is-ai-transcription" rel="noopener noreferrer"&gt;how accurate AI transcription really is&lt;/a&gt;
for what to expect.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You need a rare language or professional subtitling.&lt;/strong&gt; HappyScribe covers
150+ languages, exports 15+ formats and sells human subtitling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You're a team that lives in a shared workspace.&lt;/strong&gt; Sonix is built around
multi-seat workspaces and AI analysis of transcripts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You transcribe 80+ hours a month.&lt;/strong&gt; Rev's larger allowances (5,000 minutes
on Essentials, 10,000 on Pro) fit very high volumes; ScribeToAny Pro tops out
at 3,000 minutes a month.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;If your work is &lt;strong&gt;uploading recordings&lt;/strong&gt; — interviews, lectures, podcasts,&lt;br&gt;
videos — and you transcribe somewhere between a few hours and 50 hours a&lt;br&gt;
month, ScribeToAny is the lowest-cost way to do it among these five: $10 a&lt;br&gt;
month covers 25 hours, $20 covers 50, with 3-hour files, ~98 languages and&lt;br&gt;
every export format. You can try it on your own files with the&lt;br&gt;
&lt;a href="https://scribetoany.com/pricing" rel="noopener noreferrer"&gt;free plan&lt;/a&gt;, no card needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;p&gt;Prices and limits as published on each vendor's pricing page on September 23,&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;EUR prices are shown as published, not converted. Yearly cost per hour
divides the yearly price by 12 months of included time.&lt;/li&gt;
&lt;/ol&gt;

&lt;ul&gt;
&lt;li&gt;Otter.ai — &lt;a href="https://otter.ai/pricing" rel="noopener noreferrer"&gt;otter.ai/pricing&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Rev — &lt;a href="https://www.rev.com/pricing" rel="noopener noreferrer"&gt;rev.com/pricing&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Sonix — &lt;a href="https://sonix.ai/pricing" rel="noopener noreferrer"&gt;sonix.ai/pricing&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;HappyScribe — &lt;a href="https://www.happyscribe.com/pricing" rel="noopener noreferrer"&gt;happyscribe.com/pricing&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;ScribeToAny — &lt;a href="https://scribetoany.com/pricing" rel="noopener noreferrer"&gt;scribetoany.com/pricing&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Otter.ai, Rev, Sonix and HappyScribe are trademarks of their respective owners.&lt;br&gt;
This comparison is independent and not endorsed by any of them. Spotted a&lt;br&gt;
number that's changed? &lt;a href="https://scribetoany.com/contact" rel="noopener noreferrer"&gt;Tell us&lt;/a&gt; and we'll update it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://scribetoany.com/blog/transcription-pricing-compared" rel="noopener noreferrer"&gt;ScribeToAny blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>ai</category>
      <category>tools</category>
      <category>comparison</category>
    </item>
    <item>
      <title>Subtitle Reading Speed (CPS): The Limits, and Why AI Subtitles Break Them</title>
      <dc:creator>Ray Mac</dc:creator>
      <pubDate>Sat, 26 Sep 2026 04:35:58 +0000</pubDate>
      <link>https://dev.to/ray_mac/subtitle-reading-speed-cps-the-limits-and-why-ai-subtitles-break-them-892</link>
      <guid>https://dev.to/ray_mac/subtitle-reading-speed-cps-the-limits-and-why-ai-subtitles-break-them-892</guid>
      <description>&lt;p&gt;A subtitle can be perfectly accurate and still fail. If a line leaves the&lt;br&gt;
screen before the viewer has finished reading it, the viewer has missed it.&lt;br&gt;
&lt;strong&gt;Reading speed&lt;/strong&gt;, measured in &lt;strong&gt;characters per second (CPS)&lt;/strong&gt;, is how&lt;br&gt;
subtitlers catch that before anyone watches the video.&lt;/p&gt;

&lt;p&gt;This post covers the limits the big style guides use and how to convert&lt;br&gt;
between CPS and words per minute. It then runs a small experiment on&lt;br&gt;
machine-made subtitles. The finding surprised us: &lt;strong&gt;most "too fast" subtitles&lt;br&gt;
from AI transcription aren't too wordy. They're timed too tightly&lt;/strong&gt;, and you&lt;br&gt;
can fix most of them without deleting a single word.&lt;/p&gt;
&lt;h2&gt;
  
  
  What CPS measures
&lt;/h2&gt;

&lt;p&gt;CPS is the number of characters in a subtitle divided by how long it stays on&lt;br&gt;
screen:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CPS = characters in the cue ÷ seconds on screen
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A 42-character line shown for 2 seconds runs at 21 CPS. The same line shown for&lt;br&gt;
2.5 seconds runs at 16.8 CPS.&lt;/p&gt;

&lt;p&gt;Two details change the result, so check them before you compare numbers&lt;br&gt;
between tools:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What counts as a character.&lt;/strong&gt; We count every character the eye reads,
including spaces and punctuation, and leave out only the line break between
two rows. Some tools skip spaces or punctuation. Their numbers come out lower
for the same subtitle.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Which duration counts.&lt;/strong&gt; The time on screen is the cue's end time minus its
start time. It is not how long the person took to say the words. That
difference is the whole story of the experiment below.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The limits in the major style guides
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Guideline&lt;/th&gt;
&lt;th&gt;Reading speed&lt;/th&gt;
&lt;th&gt;Line length&lt;/th&gt;
&lt;th&gt;Lines&lt;/th&gt;
&lt;th&gt;Duration&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Netflix, English (adult)&lt;/td&gt;
&lt;td&gt;up to 20 CPS&lt;/td&gt;
&lt;td&gt;42 characters&lt;/td&gt;
&lt;td&gt;2 max&lt;/td&gt;
&lt;td&gt;5/6 s to 7 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Netflix, English (children's)&lt;/td&gt;
&lt;td&gt;up to 17 CPS&lt;/td&gt;
&lt;td&gt;42 characters&lt;/td&gt;
&lt;td&gt;2 max&lt;/td&gt;
&lt;td&gt;5/6 s to 7 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Netflix, Simplified Chinese (adult)&lt;/td&gt;
&lt;td&gt;up to 9 CPS&lt;/td&gt;
&lt;td&gt;16 characters&lt;/td&gt;
&lt;td&gt;2 max&lt;/td&gt;
&lt;td&gt;5/6 s to 7 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Netflix, Japanese&lt;/td&gt;
&lt;td&gt;up to 4 CPS&lt;/td&gt;
&lt;td&gt;13 full-width characters&lt;/td&gt;
&lt;td&gt;2 max&lt;/td&gt;
&lt;td&gt;5/6 s to 7 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BBC&lt;/td&gt;
&lt;td&gt;160–180 words per minute&lt;/td&gt;
&lt;td&gt;about 68% of a 16:9 frame width&lt;/td&gt;
&lt;td&gt;2 (3 if nothing important is covered)&lt;/td&gt;
&lt;td&gt;about 0.3 s per word minimum&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A few things stand out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Chinese and Japanese limits are far lower&lt;/strong&gt;, because each character carries
much more meaning than a Latin letter. A Latin-script limit would never flag
anything in Chinese text. You need a separate threshold for each script.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The maximum is not a target.&lt;/strong&gt; Netflix drops to 17 CPS for children's
programs, and the BBC's 160–180 words per minute works out to about 14–16
CPS (see below). Aiming below the ceiling leaves room for viewers who read
slowly or are reading in a second language.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Minimum duration matters as much as the speed limit.&lt;/strong&gt; Netflix's 5/6 of a
second (20 frames at 24 fps) exists because a flash of text is hard to catch
even when it is short.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Converting words per minute to CPS
&lt;/h2&gt;

&lt;p&gt;Speech and subtitle guidelines often use words per minute (WPM). To convert&lt;br&gt;
between the two, you need the average characters per word, including the space&lt;br&gt;
after each word.&lt;/p&gt;

&lt;p&gt;In the English sample we measured below (447 words, 2,413 characters), that&lt;br&gt;
came to &lt;strong&gt;5.4 characters per word&lt;/strong&gt;. So, for English:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Words per minute&lt;/th&gt;
&lt;th&gt;≈ CPS&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;140&lt;/td&gt;
&lt;td&gt;12.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;160&lt;/td&gt;
&lt;td&gt;14.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;180&lt;/td&gt;
&lt;td&gt;16.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;18.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;220&lt;/td&gt;
&lt;td&gt;19.8&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So the BBC's 160–180 WPM works out to about &lt;strong&gt;14–16 CPS&lt;/strong&gt; in our counting,&lt;br&gt;
noticeably below Netflix's 20 CPS adult ceiling. Audiobooks are commonly&lt;br&gt;
recommended at 150–160 WPM; the explainer voice in our test below ran at 189. A&lt;br&gt;
word-for-word subtitle of a fast talker can therefore break a 17 CPS limit on&lt;br&gt;
text alone. But as the experiment shows, text is usually not the main cause.&lt;/p&gt;

&lt;h2&gt;
  
  
  An experiment: where AI subtitles break the limit
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The audio.&lt;/strong&gt; We used the English voiceover from our own product video: 2&lt;br&gt;
minutes 48 seconds, 447 words, one clear synthetic voice with natural pauses&lt;br&gt;
between sentences. The words take up 142 seconds of that time, and within those&lt;br&gt;
seconds the voice runs at &lt;strong&gt;189 WPM (17.0 CPS)&lt;/strong&gt;. That's a brisk but normal&lt;br&gt;
explainer-video pace.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The subtitles.&lt;/strong&gt; We transcribed it with Whisper (the open-source &lt;code&gt;base&lt;/code&gt;&lt;br&gt;
model, run locally). We then cut the result into subtitle cues with the same&lt;br&gt;
splitting rules ScribeToAny's engine uses for exported SRT files:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;at most 84 characters per cue (two lines of 42);&lt;/li&gt;
&lt;li&gt;at most 7 seconds per cue;&lt;/li&gt;
&lt;li&gt;break at the end of a sentence first, then at a clause.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each cue starts at its first word and ends at its last word.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The check.&lt;/strong&gt; We counted every cue against 17 and 20 CPS, the two Netflix&lt;br&gt;
English limits, and flagged cues shorter than 5/6 of a second. We then applied&lt;br&gt;
two standard fixes, one at a time.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;Cues&lt;/th&gt;
&lt;th&gt;Median CPS&lt;/th&gt;
&lt;th&gt;Over 17 CPS&lt;/th&gt;
&lt;th&gt;Over 20 CPS&lt;/th&gt;
&lt;th&gt;Shorter than 5/6 s&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A. Cues timed to the words&lt;/td&gt;
&lt;td&gt;69&lt;/td&gt;
&lt;td&gt;19.3&lt;/td&gt;
&lt;td&gt;47 (68%)&lt;/td&gt;
&lt;td&gt;29 (42%)&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B. + end extended into the following pause&lt;/td&gt;
&lt;td&gt;69&lt;/td&gt;
&lt;td&gt;15.5&lt;/td&gt;
&lt;td&gt;20 (29%)&lt;/td&gt;
&lt;td&gt;9 (13%)&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C. + short neighbours merged&lt;/td&gt;
&lt;td&gt;41&lt;/td&gt;
&lt;td&gt;15.5&lt;/td&gt;
&lt;td&gt;10 (24%)&lt;/td&gt;
&lt;td&gt;3 (7%)&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  What the numbers say
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Timed to the words, two thirds of the cues were "too fast."&lt;/strong&gt; Yet the audio&lt;br&gt;
averages exactly 17.0 CPS while someone is talking. The average cue was not&lt;br&gt;
too wordy. Cues that start and end exactly with the speech leave no reading&lt;br&gt;
time, and the short ones suffer most. For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Need more?" was on screen for 0.36 s, which works out to 27.8 CPS.&lt;/li&gt;
&lt;li&gt;"Fix it right there." lasted 0.6 s, which works out to 31.7 CPS.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Using the pauses fixed more than half of the failures.&lt;/strong&gt; For step B, each&lt;br&gt;
cue's end time moves forward into the silence before the next line starts. We&lt;br&gt;
left an 80 ms gap (about 2 frames at 25 fps) and capped each cue at 7 seconds.&lt;br&gt;
That alone took the cues over 17 CPS from 47 to 20, without changing a word.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Merging short neighbours fixed most of the rest.&lt;/strong&gt; For step C, a cue that&lt;br&gt;
was still too fast merged with the next one when the result stayed within 84&lt;br&gt;
characters and 7 seconds. We only merged where the combined text ended at&lt;br&gt;
punctuation, so no subtitle ends mid-phrase.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A 0.44-second "Turn on" fragment rejoined the rest of its sentence, "detect
speakers, and every line gets a speaker label." The merged cue runs at
15.7 CPS.&lt;/li&gt;
&lt;li&gt;Cues over 20 CPS fell to 3 of 41.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What's left is a small editing job.&lt;/strong&gt; To bring every remaining cue to 17 CPS,&lt;br&gt;
you would need to cut &lt;strong&gt;40 characters out of 2,399, about 1.7% of the text&lt;/strong&gt;.&lt;br&gt;
To reach 20 CPS, you would cut 5 characters. That is where a human makes&lt;br&gt;
choices: dropping a filler word, or shortening "Then take it out in the format&lt;br&gt;
you need" to "Then export it."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The caveats.&lt;/strong&gt; This is one recording, with one voice, and with clean pauses&lt;br&gt;
between sentences. A real conversation has fewer pauses, overlapping speakers&lt;br&gt;
and faster bursts, so steps B and C buy less there and condensing matters&lt;br&gt;
more. But the ordering holds in general: &lt;strong&gt;fix the timing before you cut&lt;br&gt;
words.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Fixing a fast subtitle, in order
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Extend the end time into silence.&lt;/strong&gt; If the next line doesn't start
immediately, let the current cue stay on screen longer. Keep a small gap of
about 2 frames before the next cue, so the viewer can see that the subtitle
changed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Respect the minimum duration.&lt;/strong&gt; Nothing should flash on screen for less
than about 5/6 of a second, even a one-word "Yes."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Merge short neighbours.&lt;/strong&gt; Two quick lines from the same speaker often
read better as one two-line subtitle. Stay within your line and duration
limits, and don't merge across a change of speaker.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Only then, condense.&lt;/strong&gt; Remove fillers ("you know," "basically"), false
starts and repeated words. Keep the meaning, names and numbers. For a
verbatim or legal transcript you wouldn't do this, but subtitles are a
reading aid, not a record.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Break lines in sensible places.&lt;/strong&gt; When a subtitle needs two lines, break
after punctuation, or before a conjunction or preposition. Netflix's style
guide asks you not to split an article from its noun or a first name from a
last name. It also prefers a bottom-heavy shape, with the longer line
underneath.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Chinese, Japanese and Korean
&lt;/h2&gt;

&lt;p&gt;For these scripts, character counts behave differently:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Netflix's Simplified Chinese guide allows &lt;strong&gt;9 CPS and 16 characters a line&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Its Japanese guide allows &lt;strong&gt;4 CPS and 13 full-width characters&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A Latin 17 CPS threshold would pass almost any Chinese subtitle, however&lt;br&gt;
cramped. Choose the threshold by the script on screen, not by the source&lt;br&gt;
language. An English translation of a Chinese video is checked as English.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checking reading speed in ScribeToAny
&lt;/h2&gt;

&lt;p&gt;ScribeToAny's transcript editor has a &lt;strong&gt;QC mode&lt;/strong&gt; on every plan, free&lt;br&gt;
included. Turn on &lt;strong&gt;Reading-speed &amp;amp; line checks&lt;/strong&gt; and each cue shows the&lt;br&gt;
following:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;its CPS, amber when it gets close to your limit and red when it goes over;&lt;/li&gt;
&lt;li&gt;the character count of each line, with lines over the limit flagged;&lt;/li&gt;
&lt;li&gt;whether a cue is under 5/6 of a second or over 7 seconds, or too close to
the previous cue.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The defaults are the conservative ones:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;17 CPS and 42 characters a line&lt;/strong&gt; for Latin scripts;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;9 CPS and 16 characters&lt;/strong&gt; for Chinese, Japanese and Korean, picked
automatically from the text on screen.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can change both numbers, for example to 20 CPS for Netflix-style adult&lt;br&gt;
English. You can edit start and end times directly, split a cue or merge it&lt;br&gt;
with the next one, and roll back to an earlier version. Then export SRT or VTT&lt;br&gt;
with the same timings.&lt;/p&gt;

&lt;p&gt;To be clear about what it doesn't do: &lt;strong&gt;QC mode flags problems. It doesn't&lt;br&gt;
retime cues for you.&lt;/strong&gt; Steps 1 to 4 above are still your call. That's&lt;br&gt;
deliberate. A tool that silently stretches or rewrites subtitles makes&lt;br&gt;
decisions only a person watching the video can make.&lt;/p&gt;

&lt;p&gt;If you only need to adjust an existing SRT file, the free&lt;br&gt;
&lt;a href="https://scribetoany.com/tools/srt-editor" rel="noopener noreferrer"&gt;SRT editor&lt;/a&gt; runs in your browser without an account. It&lt;br&gt;
doesn't show reading-speed checks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Related:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://scribetoany.com/blog/captions-vs-subtitles" rel="noopener noreferrer"&gt;Captions vs. subtitles&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://scribetoany.com/blog/subtitle-and-transcript-file-formats" rel="noopener noreferrer"&gt;Subtitle and transcript file formats&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://scribetoany.com/blog/translate-subtitles-step-by-step" rel="noopener noreferrer"&gt;How to translate subtitles step by step&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Netflix, &lt;a href="https://partnerhelp.netflixstudios.com/hc/en-us/articles/217350977-English-USA-Timed-Text-Style-Guide" rel="noopener noreferrer"&gt;English (USA) Timed Text Style Guide&lt;/a&gt;:
42 characters a line, up to 20 CPS for adult programs and 17 for children's,
and line-break rules.&lt;/li&gt;
&lt;li&gt;Netflix, &lt;a href="https://partnerhelp.netflixstudios.com/hc/en-us/articles/215758617-Timed-Text-Style-Guide-General-Requirements" rel="noopener noreferrer"&gt;Timed Text Style Guide: General Requirements&lt;/a&gt;:
5/6-second minimum, 7-second maximum, 2 lines.&lt;/li&gt;
&lt;li&gt;Netflix, &lt;a href="https://partnerhelp.netflixstudios.com/hc/en-us/articles/215986007-Chinese-Simplified-Timed-Text-Style-Guide" rel="noopener noreferrer"&gt;Chinese (Simplified) Timed Text Style Guide&lt;/a&gt;
and &lt;a href="https://partnerhelp.netflixstudios.com/hc/en-us/articles/215767517-Japanese-Timed-Text-Style-Guide" rel="noopener noreferrer"&gt;Japanese Timed Text Style Guide&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Wikipedia, &lt;a href="https://en.wikipedia.org/wiki/Words_per_minute" rel="noopener noreferrer"&gt;Words per minute&lt;/a&gt;:
150–160 WPM recommended for audiobooks.&lt;/li&gt;
&lt;li&gt;BBC, &lt;a href="https://www.bbc.co.uk/accessibility/forproducts/guides/subtitles/" rel="noopener noreferrer"&gt;Subtitle Guidelines&lt;/a&gt;:
160–180 words per minute, line width and line count.&lt;/li&gt;
&lt;li&gt;Experiment: our own measurement on September 24, 2026, as described above.
The cue-splitting rules are the ones our transcription engine uses. Our
production service runs larger Whisper models than the &lt;code&gt;base&lt;/code&gt; model used
here, which mainly affects word accuracy rather than timing.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://scribetoany.com/blog/subtitle-reading-speed-cps" rel="noopener noreferrer"&gt;ScribeToAny blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>subtitles</category>
      <category>whisper</category>
      <category>video</category>
      <category>a11y</category>
    </item>
    <item>
      <title>How We Optimized ScribeToAny: From 3.5s Cloudflare Cold Starts to a 95+ Lighthouse Score</title>
      <dc:creator>Ray Mac</dc:creator>
      <pubDate>Tue, 08 Sep 2026 04:59:18 +0000</pubDate>
      <link>https://dev.to/ray_mac/how-we-optimized-scribetoany-from-35s-cloudflare-cold-starts-to-a-95-lighthouse-score-be</link>
      <guid>https://dev.to/ray_mac/how-we-optimized-scribetoany-from-35s-cloudflare-cold-starts-to-a-95-lighthouse-score-be</guid>
      <description>&lt;p&gt;&lt;a href="https://scribetoany.com" rel="noopener noreferrer"&gt;ScribeToAny&lt;/a&gt; is a full-stack audio and video transcription platform built on React 19 and deployed to Cloudflare Workers at the edge. Heavy GPU workloads (Whisper transcription, diarization, translation) run asynchronously on &lt;a href="https://modal.com" rel="noopener noreferrer"&gt;Modal&lt;/a&gt;, while our web application, authentication, database queries, and SSR are powered by Cloudflare Workers.&lt;/p&gt;

&lt;p&gt;While our GPU pipeline was fast and asynchronous, our front-end performance audit delivered a harsh wake-up call: real-world user monitoring (Cloudflare Observatory) showed our &lt;strong&gt;75th-percentile TTFB was 3,128 ms&lt;/strong&gt; — with over 57% of hits rated "poor". Direct &lt;code&gt;curl&lt;/code&gt; tests hitting cold edge nodes clocked TTFB between &lt;strong&gt;3.4s and 3.6s&lt;/strong&gt; for the homepage and SEO tool pages.&lt;/p&gt;

&lt;p&gt;Yet, Cloudflare's internal Worker metrics showed a median CPU wall time of &lt;strong&gt;just 5 ms&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;How could a Worker that finishes its work in 5 milliseconds take 3.5 seconds to return a response?&lt;/p&gt;

&lt;p&gt;Here is the complete engineering breakdown of how we diagnosed the cold-start bottleneck, implemented zero-staleness edge HTML caching, decoupled our client bundles, slashed mobile hydration TBT, and took ScribeToAny to a &lt;strong&gt;95+ Lighthouse performance score&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. The Paradox: Why 5ms CPU Took 3.5s TTFB
&lt;/h2&gt;

&lt;p&gt;Cloudflare Worker CPU metrics only record active handler execution. They do &lt;strong&gt;not&lt;/strong&gt; account for isolate creation, bundle downloading, and V8 script compilation.&lt;/p&gt;

&lt;p&gt;When we inspected our deployment pipeline, two factors collided:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;An 11MB Worker Bundle&lt;/strong&gt;: Because ScribeToAny includes rich SEO tool routes (over 80 &lt;a href="https://scribetoany.com/tools" rel="noopener noreferrer"&gt;audio/video conversion and transcription tools&lt;/a&gt;), markdown renderers, and format converter utilities, the bundled worker script reached ~11MB.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Low Baseline Traffic Density (~0.1 req/s)&lt;/strong&gt;: With sparse initial traffic, Cloudflare edge PoPs frequently evict idle V8 isolates.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Almost every new visitor — especially organic search traffic arriving from Google — hit a brand-new, cold isolate. Before running our 5ms SSR handler, V8 had to load and compile an 11MB JavaScript payload. The result was a 3-second cold-start penalty on first visit, completely undermining the user experience and SEO ranking potential.&lt;/p&gt;

&lt;p&gt;Furthermore, on mobile devices, initial Lighthouse runs flagged heavy Total Blocking Time (TBT): third-party authentication scripts and oversized vendor chunks were monopolizing the main thread during hydration.&lt;/p&gt;

&lt;p&gt;We solved this through a systematic four-layer optimization strategy.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Layer 1: Edge HTML Caching with Zero-Staleness Build Invalidation
&lt;/h2&gt;

&lt;p&gt;Cloudflare Workers do not automatically cache dynamic SSR responses. Every incoming &lt;code&gt;GET&lt;/code&gt; request was hitting our Worker, forcing a cold-start compilation.&lt;/p&gt;

&lt;p&gt;Because our marketing pages, blog, and SEO tools are public and identical for anonymous visitors of the same URL, we implemented edge-level HTML caching directly in &lt;code&gt;src/server.ts&lt;/code&gt; using the Workers Cache API (&lt;code&gt;caches.default&lt;/code&gt;).&lt;/p&gt;

&lt;h3&gt;
  
  
  Strict Isolation &amp;amp; Bypass Rules
&lt;/h3&gt;

&lt;p&gt;Edge caching dynamic web apps requires strict safety boundaries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Session isolation&lt;/strong&gt;: If the incoming request has a &lt;code&gt;better-auth.session_token&lt;/code&gt; cookie, cache reading and writing are completely bypassed. Logged-in users always receive fresh, personalized SSR.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Strict allowlist&lt;/strong&gt;: Caching is restricted strictly to anonymous public pages (&lt;code&gt;/&lt;/code&gt;, &lt;code&gt;/pricing&lt;/code&gt;, &lt;code&gt;/about&lt;/code&gt;, &lt;code&gt;/changelog&lt;/code&gt;, &lt;code&gt;/tools/*&lt;/code&gt;, &lt;code&gt;/blog/*&lt;/code&gt;, and legal pages). Dynamic routes (&lt;code&gt;/dashboard&lt;/code&gt;, &lt;code&gt;/api/*&lt;/code&gt;, &lt;code&gt;/settings&lt;/code&gt;) are never cached.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cache write condition&lt;/strong&gt;: Only responses with HTTP status &lt;code&gt;200&lt;/code&gt;, &lt;code&gt;Content-Type: text/html&lt;/code&gt;, and no &lt;code&gt;Set-Cookie&lt;/code&gt; header are stored.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// src/server.ts (abridged)&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;canCache&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;method&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;GET&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt;
  &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nf"&gt;hasSessionCookie&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt;
  &lt;span class="nf"&gt;isCacheablePath&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;deLocalizeUrl&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;URL&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nx"&gt;pathname&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cache&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;caches&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;unknown&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;default&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Cache&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="k"&gt;default&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cacheKey&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;canCache&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nf"&gt;edgeCacheKey&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;canCache&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;hit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;match&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cacheKey&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;hit&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;headers&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Headers&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;hit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;X-Edge-Cache&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;HIT&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;hit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;hit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;headers&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Solving the Stale Content Problem
&lt;/h3&gt;

&lt;p&gt;The biggest danger of &lt;code&gt;caches.default&lt;/code&gt; on Cloudflare Workers is that &lt;strong&gt;a new Worker deployment does not purge the cache&lt;/strong&gt;. With a standard URL key, publishing a new blog post or fixing a bug would leave stale HTML served to users for days.&lt;/p&gt;

&lt;p&gt;We solved this by injecting a compile-time build identifier into the cache key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// vite.config.ts&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="nf"&gt;defineConfig&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;define&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;__EDGE_BUILD_ID__&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;36&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="c1"&gt;// ...&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In &lt;code&gt;src/server.ts&lt;/code&gt;, we construct an internal cache key request tagged with &lt;code&gt;__EDGE_BUILD_ID__&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;edgeCacheKey&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nx"&gt;Request&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;URL&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;searchParams&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;__ev&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;__EDGE_BUILD_ID__&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This internal query parameter is only passed to &lt;code&gt;cache.match&lt;/code&gt; and &lt;code&gt;cache.put&lt;/code&gt;; it is never exposed to the client or upstream origin.&lt;/p&gt;

&lt;p&gt;When a new version deploys, &lt;code&gt;__EDGE_BUILD_ID__&lt;/code&gt; changes. Old cache entries become instantly unreachable and expire naturally on their TTL, while the new release begins populating immediately. We can safely set &lt;code&gt;s-maxage=86400&lt;/code&gt; (24h) and &lt;code&gt;stale-while-revalidate=604800&lt;/code&gt; (7 days) without ever worrying about stale HTML after a release.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Result&lt;/strong&gt;: Edge cache hits now return in &lt;strong&gt;under 45 milliseconds&lt;/strong&gt; directly from the nearest Cloudflare edge PoP, completely bypassing Worker isolate cold starts.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Layer 2: Main Bundle Decoupling &amp;amp; Vendor Splitting
&lt;/h2&gt;

&lt;p&gt;Serving HTML quickly is only half the battle; if the browser has to parse hundreds of kilobytes of unoptimized JavaScript before becoming interactive, the experience still stumbles.&lt;/p&gt;

&lt;h3&gt;
  
  
  Decoupling Global Configs
&lt;/h3&gt;

&lt;p&gt;Our site configuration (&lt;code&gt;src/config/website.ts&lt;/code&gt;) originally bundled metadata, navigation menus, multilingual pricing matrices, and 84 tool route names into one giant object. Because the root layout imported &lt;code&gt;websiteConfig&lt;/code&gt;, every visitor downloaded definitions for 84 tools they hadn't visited.&lt;/p&gt;

&lt;p&gt;We extracted tool metadata into &lt;code&gt;src/config/navbar-tool-names.ts&lt;/code&gt; and pricing calculations into &lt;code&gt;src/config/price-plans.ts&lt;/code&gt;, ensuring tool pages only load their own definitions on demand.&lt;/p&gt;

&lt;h3&gt;
  
  
  Vite Manual Chunks
&lt;/h3&gt;

&lt;p&gt;By default, Vite bundled React, TanStack Query, and Zod into monolithic client bundles. In &lt;code&gt;vite.config.ts&lt;/code&gt;, we configured explicit vendor chunks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// vite.config.ts&lt;/span&gt;
&lt;span class="nx"&gt;build&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;rollupOptions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;output&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nf"&gt;manualChunks&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@tabler/icons-react&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;vendor-icons&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
          &lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node_modules/react/&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
          &lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node_modules/react-dom/&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
          &lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node_modules/scheduler/&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;vendor-react&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
          &lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node_modules/@tanstack/react-query&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
          &lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node_modules/@tanstack/query-core&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;vendor-query&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node_modules/zod/&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;vendor-zod&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This isolates shared dependencies, improving long-term browser caching across route transitions and preventing small code changes from invalidating vendor libraries.&lt;/p&gt;

&lt;h3&gt;
  
  
  Scoping Global Providers
&lt;/h3&gt;

&lt;p&gt;In &lt;code&gt;src/routes/__root.tsx&lt;/code&gt;, Radix &lt;code&gt;TooltipProvider&lt;/code&gt; originally wrapped the entire application tree. This forced React to initialize tooltip context on every public landing page even though tooltips were only used inside the logged-in dashboard and transcription editor. We removed &lt;code&gt;TooltipProvider&lt;/code&gt; from the root route and scoped it exclusively to the editor and dashboard components.&lt;/p&gt;

&lt;p&gt;Similarly, we extracted &lt;code&gt;prose.css&lt;/code&gt; from the global &lt;code&gt;styles.css&lt;/code&gt;. Markdown typography styling is now loaded strictly on blog, legal, and documentation routes, eliminating unused CSS overhead on marketing pages.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Layer 3: Offscreen Lazy-Loading &amp;amp; Slashing Hydration TBT
&lt;/h2&gt;

&lt;p&gt;On mobile devices, client-side hydration was choking the CPU. When a browser downloads a page, hydrating every single DOM node on a long marketing page blocks user taps and scrolls (high Total Blocking Time).&lt;/p&gt;

&lt;h3&gt;
  
  
  Synchronous Above-the-Fold, Lazy Below-the-Fold
&lt;/h3&gt;

&lt;p&gt;In &lt;code&gt;src/components/blocks/homepage.tsx&lt;/code&gt;, we split the landing page into two categories:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Above-the-Fold (Synchronous)&lt;/strong&gt;: &lt;code&gt;HeroSection&lt;/code&gt;, &lt;code&gt;TrustStrip&lt;/code&gt;, and &lt;code&gt;WhisperTechSection&lt;/code&gt; are imported synchronously. They render immediately during SSR and hydrate on frame 1 for instant First Contentful Paint (FCP).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Below-the-Fold (Lazy Loaded)&lt;/strong&gt;: &lt;code&gt;FeaturesSection&lt;/code&gt;, &lt;code&gt;Features2Section&lt;/code&gt;, &lt;code&gt;StatsSection&lt;/code&gt;, &lt;code&gt;CallToActionSection&lt;/code&gt;, &lt;code&gt;PricingSection&lt;/code&gt;, &lt;code&gt;FaqSection&lt;/code&gt;, and &lt;code&gt;NewsletterCard&lt;/code&gt; are wrapped in &lt;code&gt;React.lazy()&lt;/code&gt; and &lt;code&gt;Suspense&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;To prevent Cumulative Layout Shift (CLS) when these components resolve, each &lt;code&gt;Suspense&lt;/code&gt; boundary has an explicit minimum-height skeleton fallback matching the component's rendered height:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="c1"&gt;// src/components/blocks/homepage.tsx&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;HomePage&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt; &lt;span class="na"&gt;className&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"flex flex-col"&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;HeroSection&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;TrustStrip&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;WhisperTechSection&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Suspense&lt;/span&gt; &lt;span class="na"&gt;fallback&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt; &lt;span class="na"&gt;className&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"min-h-[650px]"&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
        &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;FeaturesSection&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nc"&gt;Suspense&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Suspense&lt;/span&gt; &lt;span class="na"&gt;fallback&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt; &lt;span class="na"&gt;className&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"min-h-[550px]"&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
        &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Features2Section&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nc"&gt;Suspense&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Suspense&lt;/span&gt; &lt;span class="na"&gt;fallback&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt; &lt;span class="na"&gt;className&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"min-h-[300px]"&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
        &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;CallToActionSection&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nc"&gt;Suspense&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="cm"&gt;/* ... */&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This guaranteed &lt;strong&gt;0.00 CLS&lt;/strong&gt; while cutting the initial hydration JavaScript payload by over 40%.&lt;/p&gt;

&lt;h3&gt;
  
  
  Deferring Google One Tap to Idle
&lt;/h3&gt;

&lt;p&gt;One of the largest contributors to mobile Total Blocking Time was Google One Tap (&lt;code&gt;authClient.oneTap()&lt;/code&gt;). Previously, the Google Identity Services script executed during initial React hydration, spinning up iframe bridges and network checks while the user was trying to interact with the hero section.&lt;/p&gt;

&lt;p&gt;We moved Google One Tap out of the critical rendering path by wrapping it in &lt;code&gt;requestIdleCallback&lt;/code&gt; (with an 8-second fallback timeout) and attaching one-time event listeners to user interactions (&lt;code&gt;pointerdown&lt;/code&gt;, &lt;code&gt;touchstart&lt;/code&gt;, &lt;code&gt;scroll&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="c1"&gt;// src/routes/__root.tsx (abridged)&lt;/span&gt;
&lt;span class="nf"&gt;useEffect&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;isOneTapEnabled&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;isPending&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;session&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;executed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;trigger&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;executed&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nx"&gt;executed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nf"&gt;cleanup&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;void&lt;/span&gt; &lt;span class="nx"&gt;authClient&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;oneTap&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;interactionEvents&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;pointerdown&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;touchstart&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;scroll&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;evt&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;interactionEvents&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;addEventListener&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;evt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;trigger&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;once&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;passive&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;requestIdleCallback&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;function&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;requestIdleCallback&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;trigger&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;timeout&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;8000&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nf"&gt;setTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;trigger&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;7000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="c1"&gt;// ...&lt;/span&gt;
&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;isPending&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;session&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The result: zero main-thread interference during initial paint.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Layer 4: LCP &amp;amp; Accessibility Polish
&lt;/h2&gt;

&lt;p&gt;With TTFB and hydration fixed, we addressed visual rendering timing and audit metrics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Critical CSS Preloading&lt;/strong&gt;: Added &lt;code&gt;&amp;lt;link rel="preload" as="style" href={appCss} /&amp;gt;&lt;/code&gt; in &lt;code&gt;src/routes/__root.tsx&lt;/code&gt; to eliminate stylesheet render blocking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hero H1 Animation Delay Removal&lt;/strong&gt;: In &lt;code&gt;src/components/blocks/hero.tsx&lt;/code&gt;, our primary H1 heading previously had a subtle entrance fade-in animation delay. Removing the delay allowed Lighthouse to record Largest Contentful Paint (LCP) the instant the first paint completed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accessibility &amp;amp; Contrast&lt;/strong&gt;: Fixed &lt;code&gt;aria-orientation="horizontal"&lt;/code&gt; on button toggle groups, added descriptive &lt;code&gt;aria-label&lt;/code&gt; tags, and increased primary button color contrast to meet WCAG AA standards.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Descriptive Anchor Text&lt;/strong&gt;: Replaced ambiguous "Learn more" link texts with descriptive link destinations ("Security →"), satisfying Lighthouse SEO crawlability checks.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  6. The Results: Lighthouse 95+ &amp;amp; Core Web Vitals
&lt;/h2&gt;

&lt;p&gt;After deploying these four layers, we ran Google PageSpeed Insights on &lt;a href="https://scribetoany.com" rel="noopener noreferrer"&gt;scribetoany.com&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Desktop Performance: 95 / 100
&lt;/h3&gt;

&lt;p&gt;On desktop devices, our score surged to &lt;strong&gt;95 Performance&lt;/strong&gt;, &lt;strong&gt;100 Accessibility&lt;/strong&gt;, &lt;strong&gt;96 Best Practices&lt;/strong&gt;, and &lt;strong&gt;100 SEO&lt;/strong&gt;, with a 3/3 score on Agentic Browsing audits.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftpq7113fulc8bjy4poxi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftpq7113fulc8bjy4poxi.png" alt="Google PageSpeed Insights Desktop Report for ScribeToAny" width="800" height="688"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Score / Value&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Performance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;95&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;🟢 Excellent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Accessibility&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;🟢 Perfect&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Best Practices&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;96&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;🟢 Excellent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SEO&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;🟢 Perfect&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Agentic Browsing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3 / 3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;🟢 Passed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cumulative Layout Shift (CLS)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;🟢 Zero Shift&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Mobile Performance: 78 / 100
&lt;/h3&gt;

&lt;p&gt;On simulated mobile networks with aggressive 4G throttling and restricted mobile CPU profiles, our performance jumped to &lt;strong&gt;78 Performance&lt;/strong&gt;, alongside flawless &lt;strong&gt;100 Accessibility&lt;/strong&gt; and &lt;strong&gt;100 SEO&lt;/strong&gt; scores:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1w3h08bhhpybq2j2yg64.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1w3h08bhhpybq2j2yg64.png" alt="Google PageSpeed Insights Mobile Report for ScribeToAny" width="800" height="660"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before Optimization&lt;/th&gt;
&lt;th&gt;After Optimization&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;P75 Edge TTFB&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~3,128 ms (57% Poor)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;&amp;lt; 50 ms (Edge HIT)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cold Worker TTFB&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~3,500 ms&lt;/td&gt;
&lt;td&gt;Bypassed for all anon visitors&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Desktop Performance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~68&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;95&lt;/strong&gt; 🟢&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Mobile Performance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~42&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;78&lt;/strong&gt; 🟡&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cumulative Layout Shift&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.03&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0.00&lt;/strong&gt; 🟢&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Accessibility&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;92&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;100&lt;/strong&gt; 🟢&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SEO&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;90&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;100&lt;/strong&gt; 🟢&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Key Lessons for Full-Stack Edge Applications
&lt;/h2&gt;

&lt;p&gt;Building on Cloudflare Workers and a modern full-stack React framework gives you unprecedented global reach, but edge runtimes operate under different rules than traditional Node servers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Beware the "Fast Handler, Slow Isolate" trap&lt;/strong&gt;: If your Worker takes 5ms of CPU time but bundle size is 10MB+, your users are waiting on V8 compilation, not code execution.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Edge Cache is mandatory for SSR&lt;/strong&gt;: You don't need a static site generator (SSG) to get static speeds. Using Cloudflare's &lt;code&gt;caches.default&lt;/code&gt; with build-ID key invalidation gives you static speed with all the flexibility of SSR.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guard your cache keys with build tags&lt;/strong&gt;: Never deploy edge caches without an automated invalidation strategy. Folding compile-time hashes into internal cache keys guarantees zero-downtime freshness.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hydrate lazily, paint immediately&lt;/strong&gt;: Render above-the-fold components synchronously, defer offscreen blocks with explicit height skeletons, and push heavy third-party scripts (like Google Identity Services) to &lt;code&gt;requestIdleCallback&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Speed is a feature. By aligning our edge architecture with browser execution priorities, ScribeToAny now delivers an instantaneous experience to users worldwide. &lt;/p&gt;

</description>
      <category>cloudflare</category>
      <category>webdev</category>
      <category>performance</category>
      <category>react</category>
    </item>
    <item>
      <title>Running Whisper on Modal from Cloudflare Workers</title>
      <dc:creator>Ray Mac</dc:creator>
      <pubDate>Thu, 03 Sep 2026 05:32:39 +0000</pubDate>
      <link>https://dev.to/ray_mac/running-whisper-on-modal-from-cloudflare-workers-2hla</link>
      <guid>https://dev.to/ray_mac/running-whisper-on-modal-from-cloudflare-workers-2hla</guid>
      <description>&lt;p&gt;A Cloudflare Worker can't run a multi-minute GPU job, and it can't open the gRPC connection Modal's Python SDK uses to &lt;code&gt;spawn()&lt;/code&gt; one. Two hard "no"s — and yet the transcription in &lt;a href="https://scribetoany.com" rel="noopener noreferrer"&gt;ScribeToAny&lt;/a&gt; runs on Modal GPUs while the whole web app runs on Workers. The trick is to stop treating Modal as an SDK and start treating it as an HTTP endpoint you fire-and-forget, with the GPU box calling back over a signed webhook. Here's the whole design, with the edges that actually bit.&lt;/p&gt;

&lt;p&gt;The web app runs entirely on Cloudflare Workers; Whisper (plus an optional translation pass) runs on GPUs on &lt;a href="https://modal.com" rel="noopener noreferrer"&gt;Modal&lt;/a&gt;. Getting two runtimes with opposite shapes to cooperate was the most interesting part of the build — because the obvious way is impossible on Workers.&lt;/p&gt;

&lt;h2&gt;
  
  
  The constraint
&lt;/h2&gt;

&lt;p&gt;A Worker is not a server. It wakes up on a request, gets a small CPU budget, and is expected to return quickly. It has no long-lived process to babysit a job that takes minutes, and it can't open the gRPC connection Modal's Python SDK uses to call &lt;code&gt;Function.spawn()&lt;/code&gt;. Two non-starters, same conclusion:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You cannot run the transcription &lt;em&gt;in&lt;/em&gt; the Worker. A ten-minute podcast is not a request-scoped workload.&lt;/li&gt;
&lt;li&gt;You cannot even use Modal's normal client to &lt;em&gt;start&lt;/em&gt; the job. There's no gRPC, no Python, no persistent socket.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The naive version — "await the transcription and return the transcript" — dies on the first point. So the design has to be asynchronous from the very first line, and the Worker's entire job shrinks to a handful of sub-second HTTP calls.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of the answer
&lt;/h2&gt;

&lt;p&gt;Treat Modal not as an SDK but as an HTTP endpoint. Modal lets you expose a web endpoint that, when hit, &lt;em&gt;spawns&lt;/em&gt; the real GPU function and returns immediately. So the flow becomes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The browser uploads the media straight to R2 (never through the Worker).&lt;/li&gt;
&lt;li&gt;The Worker presigns a read URL, writes a &lt;code&gt;queued&lt;/code&gt; job row, and fires one POST at Modal's web endpoint. Modal acks with a call id and starts the GPU work in the background.&lt;/li&gt;
&lt;li&gt;When the engine finishes (or fails, or just wants to report progress), it POSTs back to a webhook on the Worker, signed with a shared secret.&lt;/li&gt;
&lt;li&gt;The frontend polls the job row and lights up when it flips to &lt;code&gt;done&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The Worker only ever does three quick things: presign, spawn, apply-callback. None of them wait on a GPU. That's the whole trick — and the rest of the work is making it survive the real world, where webhooks get lost and callbacks arrive twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Uploading without touching the Worker
&lt;/h2&gt;

&lt;p&gt;Media never streams through the Worker — that would blow the CPU budget and buy nothing. The browser gets a presigned R2 &lt;code&gt;PUT&lt;/code&gt; and uploads directly. We issue it &lt;em&gt;intent-first&lt;/em&gt;: a &lt;code&gt;pending&lt;/code&gt; row is written &lt;strong&gt;before&lt;/strong&gt; the URL is signed, so an upload that's abandoned mid-flight still leaves a trace we can sweep later.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// createUploadUrl (server function) — abridged&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;insert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;userFiles&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;values&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;r2Key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;pending&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="cm"&gt;/* … */&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;uploadUrl&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;presignR2Url&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r2Key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;PUT&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;PRESIGN_PUT_TTL&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;fileId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;uploadUrl&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After the &lt;code&gt;PUT&lt;/code&gt; succeeds the client calls &lt;code&gt;finalizeUpload&lt;/code&gt;, which flips &lt;code&gt;pending → uploaded&lt;/code&gt; and corrects the size from R2's &lt;code&gt;HEAD&lt;/code&gt; (never trust a client-reported byte count). Rows that never reach &lt;code&gt;uploaded&lt;/code&gt; are garbage-collected by a cron sweep — more on that below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Firing the job
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;transcribeFile&lt;/code&gt; is where the async handoff happens. It checks quota, presigns a &lt;strong&gt;GET&lt;/strong&gt; so the engine can read the audio back out of R2, inserts a &lt;code&gt;queued&lt;/code&gt; job behind an atomic concurrency guard, and spawns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;audioUrl&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;presignR2Url&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;r2Key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;GET&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;PRESIGN_GET_TTL&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="c1"&gt;// … insert the queued job row (atomic guard) …&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;spawnTranscription&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;jobId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;audioUrl&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="cm"&gt;/* mode, language, targetLang … */&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;jobId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;queued&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The spawn itself is deliberately dumb: one &lt;code&gt;fetch&lt;/code&gt;, and it resolves the moment Modal &lt;em&gt;accepts&lt;/em&gt; the job — not when transcription finishes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;MODAL_TRANSCRIBE_URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="c1"&gt;// Modal Proxy Auth — token id/secret, not sent in the body&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Modal-Key&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;MODAL_KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Modal-Secret&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;MODAL_SECRET&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;job_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;jobId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;callback_url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;APP_URL&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/api/transcripts/webhook`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;audio_url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;audioUrl&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;model_size&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;beam_size&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;language&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;target_lang&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;targetLang&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// null ⇒ no translation leg at all&lt;/span&gt;
  &lt;span class="p"&gt;}),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Failed to start transcription (&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;)`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things worth calling out. The callback URL is handed to the engine in the request, so the engine never has to know our topology. And the signing secret is &lt;strong&gt;pre-shared&lt;/strong&gt; (a Modal secret that equals our &lt;code&gt;MODAL_WEBHOOK_SECRET&lt;/code&gt;) — it is never put in a request body in either direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  The callback: verify, then apply idempotently
&lt;/h2&gt;

&lt;p&gt;Everything interesting now happens in the webhook. First, authenticate it. The engine signs the &lt;strong&gt;raw body&lt;/strong&gt; with HMAC-SHA256 and sends the hex digest in &lt;code&gt;X-Webhook-Signature&lt;/code&gt;. On Workers there's no Node &lt;code&gt;crypto&lt;/code&gt;, so this is WebCrypto, and the comparison is timing-safe:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;verifyWebhookSignature&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;rawBody&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;signature&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;secret&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;signature&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;crypto&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;subtle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;importKey&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;raw&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;TextEncoder&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;secret&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;HMAC&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;hash&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;SHA-256&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;sign&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sig&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;crypto&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;subtle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sign&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;HMAC&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;TextEncoder&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;rawBody&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;expected&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[...&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Uint8Array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sig&lt;/span&gt;&lt;span class="p"&gt;)].&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;padStart&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;0&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;timingSafeEqual&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;expected&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;signature&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/^sha256=/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toLowerCase&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You have to hash the &lt;em&gt;raw&lt;/em&gt; bytes, not a re-serialized object — &lt;code&gt;JSON.parse&lt;/code&gt; then &lt;code&gt;JSON.stringify&lt;/code&gt; will reorder keys and change whitespace, and your signature will never match. So the handler reads &lt;code&gt;await request.text()&lt;/code&gt; and verifies before it parses.&lt;/p&gt;

&lt;p&gt;Then apply the result. The single most important property here is &lt;strong&gt;idempotency&lt;/strong&gt;, because a webhook you don't 200 fast enough gets retried, and a retried callback must not double-apply. The rule is one line: once a job is terminal, ignore repeats.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;done&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;failed&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;applied&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt; &lt;span class="c1"&gt;// already terminal — no-op&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The HTTP status codes are chosen to steer the engine's retry behaviour:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Response&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Bad/missing signature&lt;/td&gt;
&lt;td&gt;&lt;code&gt;401&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Reject outright&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unknown &lt;code&gt;job_id&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;200&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Ack so the engine &lt;strong&gt;stops&lt;/strong&gt; retrying a job we'll never have&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Applied (or already terminal)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;200&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Success&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Our own DB threw&lt;/td&gt;
&lt;td&gt;&lt;code&gt;500&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Ask the engine to &lt;strong&gt;retry&lt;/strong&gt; — the callback was valid, we just fumbled it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That third row is the subtle one. An unknown job isn't an error to bubble up; it's a dead letter, and the kindest thing you can do is acknowledge it so the sender gives up.&lt;/p&gt;

&lt;h2&gt;
  
  
  The sharp edges
&lt;/h2&gt;

&lt;p&gt;The happy path above is maybe a third of the code. The rest is everything that goes wrong when one side of an async contract can vanish.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lost webhooks.&lt;/strong&gt; If the engine crashes, or the callback is dropped, the job sits in &lt;code&gt;transcribing&lt;/code&gt; forever. So a cron job reconciles: any job that's been quiet past a timeout is marked &lt;code&gt;failed&lt;/code&gt;, and the same tick sweeps orphaned R2 uploads that never finalized. Workers cron triggers are perfect for this — it's the backstop that makes the optimistic async path safe to rely on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Never kill a healthy job by accident.&lt;/strong&gt; We also have a liveness probe that can ask Modal whether a call is still running. It returns a deliberately three-valued answer — &lt;code&gt;running&lt;/code&gt;, a terminal state, or &lt;code&gt;null&lt;/code&gt; meaning &lt;em&gt;"don't know"&lt;/em&gt; (probe not configured, request failed, unparseable). Callers must treat &lt;code&gt;null&lt;/code&gt; as &lt;em&gt;no information&lt;/em&gt; and leave the job alone. Collapsing "I couldn't reach the probe" into "the job is dead" would have the reconciler executing healthy jobs the moment the probe has a bad minute.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Segments live in R2, not the database.&lt;/strong&gt; A terminal callback does &lt;strong&gt;not&lt;/strong&gt; ship the transcript inline. The engine writes segments to R2 as &lt;code&gt;{job_id}.tsv&lt;/code&gt;; the &lt;code&gt;done&lt;/code&gt; webhook carries only metadata (language, duration, cost, timing). The app reads the TSV on demand and generates SRT/VTT/TXT/PDF from it. Keeping thousands of cue rows out of D1 keeps the callback small and the job table narrow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two clocks in one job.&lt;/strong&gt; Add the optional translation pass and a single job now finishes two things at different times. The transcript can be &lt;code&gt;done&lt;/code&gt; while the translation is still running. So the translation update is applied &lt;em&gt;before&lt;/em&gt; the terminal-state early-return — otherwise a translation ping arriving after the transcript finished would hit the "already terminal, no-op" branch and the translation row would be stuck at &lt;code&gt;queued&lt;/code&gt; forever. Two independent legs, one job row, and the order of those two checks is load-bearing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is actually a good fit
&lt;/h2&gt;

&lt;p&gt;It's tempting to read all this as fighting the platform. It isn't. Once the transcription is off the Worker, everything the Worker &lt;em&gt;does&lt;/em&gt; keep is exactly what the Workers model is good at: short, stateless, I/O-bound HTTP handlers with a cron backstop and a durable store (D1 + R2) holding the state between them. The GPU box does GPU work; the edge does edge work; a signed webhook and an idempotent apply are the seam.&lt;/p&gt;

&lt;p&gt;The constraint that looked fatal — "you can't run the job here" — turned out to be the thing that produced a clean design. The Worker never blocks, the job survives a dropped callback, and a retried webhook is a no-op. That's the whole system.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;ScribeToAny is built on TanStack Start + React on Cloudflare Workers (D1 + R2), with the transcription engine on Modal. Originally posted &lt;a href="https://scribetoany.com/blog/running-whisper-on-modal-from-cloudflare-workers" rel="noopener noreferrer"&gt;on the ScribeToAny blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cloudflare</category>
      <category>serverless</category>
      <category>webdev</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
