Jacek Migdał
Varşova, Mazovya Voyvodalığı, Polonya
7 B takipçi
500+ bağlantı
Jacek Migdał ile ortal bağlantıları görüntüle
Jacek sizi Quesma şirketindeki 10 üzerinde kişiyle tanıştırabilir
veya
LinkedIn‘de yeni misiniz? Hemen katılın
Devam Et’i tıklayarak veya oturum açarak LinkedIn Kullanıcı Anlaşması’nı, Gizlilik Politikası’nı ve Çerez Politikası’nı kabul edersiniz.
Jacek Migdał ile ortal bağlantıları görüntüle
veya
LinkedIn‘de yeni misiniz? Hemen katılın
Devam Et’i tıklayarak veya oturum açarak LinkedIn Kullanıcı Anlaşması’nı, Gizlilik Politikası’nı ve Çerez Politikası’nı kabul edersiniz.
Hakkında
Building RL environments and evaluations for frontier AI labs and enterprises. Your…
Jacek Migdał adlı kullanıcıya ait yazılar
-
Are unexpected costs the cloud equivalent of environmental pollution?
Are unexpected costs the cloud equivalent of environmental pollution?
https://jacek.migdal.
12
Faaliyet
7 B takipçi
-
Jacek Migdał bunu paylaştıThis week at Quesma: Malgorzata Piotrowska, Cezary Piwowarczyk, and Bartosz Kotrys attended the sold-out Tokenomicon + FinOps X in Amsterdam, alongside 500+ practitioners who own AI budgets. Our technical benchmarks are gaining traction. Jacek Migdał spent two days with peer founders at the Inovo.vc Magia Summit. Piotr Migdał taught the next generation of engineers at the Model Trainers workshop. VCs are knocking too. Piotr Bukanski from Acurio Ventures visited our Warsaw office. Meanwhile, we merged many PRs and dogfooded our product. Next month, we will launch our product: analytics for agentic coding and go on a US tour: SF Bay Area, Oct 19 to 30, with a booth at AGNTCon + MCPCon in San Jose on Oct 22 and 23.
-
Jacek Migdał bunu paylaştıStudents had one Friday evening, 18 GB of coding-agent transcripts, and a single question: where are we burning tokens? They investigated wasted tokens, agents stuck in loops, and whether swearing at AI changes its behavior. This was Quesma’s first token economics hackathon. Please read our blog post. Link in the comments. Congratulations to Ekipa (Paweł Juzek, Natalia Sołoducha, Maja Kowalska) for winning, with a special mention to Zespół 1 (Michał Sarzała, Piotr Z., Tomasz Święcki) for my favorite presentation. Thanks to Kevin Czupryński (STARTUP FOUNDERS STARS), Bartosz Podgórski, Tomasz Swieboda (Inovo.vc), and Malgorzata Piotrowska (Quesma) for organizing.
-
Jacek Migdał bunu paylaştıMonday: reunited with Werner Vogels, AWS CTO, 11 years after Tel Aviv (pic in high-vis vest). Now in Warsaw, with Polish founders Adam Dancewicz, Prince Canuma, Dzmitry Kamarouski, and many more. Thanks, Pavel Karatkevich, for organizing. Tuesday: Quesma changed offices and had lunch with rising star Maja Kowalska. Upcoming: Friday: Vibestar hackathon Master of Token Economics, hosted at Quesma. Sep 21 to 27: Warsaw Model Trainers at Kolektyw3. Four evening workshops, then a 46-hour hackathon with one question: can a small Polish language model pass the final high school exam? Quesma is sponsoring, and our Piotr Migdał is co-leading the training workshop. Back-to-school season is here. Sign-up links in the comments.
-
Jacek Migdał bunu paylaştıBack-to-school season at Quesma. Two opportunities in Warsaw, Poland, to embrace AI. Friday, Sep 11: We’re hosting a hackathon at our office, co-organized with STARTUP FOUNDERS STARS and Inovo.vc. Students and builders get real data, a real problem, and $100 in token credits. Your chance to become a master of the token economy. Sep 21 to 27: Warsaw Model Trainers at Kolektyw3. Four evening workshops, then a 46-hour hackathon with one question: can a small Polish language model pass the final high school exam? Quesma is sponsoring, and our Piotr Migdał is co-leading the training workshop. Links in the comments. Come and build.
-
Jacek Migdał bunu paylaştıNearly half of my August 2026 Claude Code usage came from Opus subagents, with Fable as the driver. With Fable 5.1 out and its cheaper cache reads, I'm trying a crazy setting: Fable all-in. The math behind it. I'm on the $200 Claude Max 20x plan. The token equivalent of my August usage: $5,502. That's a 27.5x subsidy (July: 15x). I promise not to sue Anthropic, unlike the folks in Kahn v. Anthropic who say Max 20x delivers too little. For scale, SemiAnalysis stress-tested Max 20x in June 2026 and hit a ceiling of about $8,000 a month. I'm at 69% of it from organic use. The same tokens repriced on OpenRouter: GPT-5 mini $203, DeepSeek V4 Flash $94. Where the money goes: cache traffic is 89% of my Claude Code cost. Cache reads alone are 64%, output tokens 11%, my own input under 0.1%. It's all about cache. Fable 5.1 drops cache reads from $1 to $0.25 per MTok, so the Opus override goes. See the Quesma blog post in the comments.
-
Jacek Migdał bunu paylaştıWrote a new blog post about Codex Pricing with experiments and a comparison to Claude Code. Seat-based subsidy is real. Thanks Piotr Migdał, Bartosz Kotrys, Przemysław Hejman and Malgorzata Piotrowska for your feedback. CC Quesma
-
Jacek Migdał bunu paylaştıOver the last 10 days, I met three great technical leaders in person. Rotem Tamir is launching AI-native CPA firms in Israel. Previously, as Co-Founder and CTO of Ariga, he built Atlas, a popular database schema tool, backed by $18M in VC funding. Tejas Khot, Staff Machine Learning Engineer at Abnormal, one of the fastest-growing cybersecurity companies of all time, valued at over $5B. He was the founding engineer of the BESP division and worked directly with the CTO. Piotr Mazurek, Member of Technical Staff at Liquid AI, behind the crazy tok/s of LFM2.5. Previously did inference at Aleph Alpha and is the author of some of the best writing on LLM inference economics. If you’re building at the frontier, happy to talk. CC Quesma
-
Jacek Migdał bunu paylaştıSame tokens, same model, up to a 40x price gap. Another Quesma blog post hit the Hacker News homepage. Priced my own Claude Code usage at API list rates: - Max 5x ($100/mo): $2,092 (21x) - Max 20x ($200/mo): $2,986 (15x) The seat plans are a subsidy. Enjoy the buffet while it's open. Link to full Math in the comments.
-
Jacek Migdał bunu paylaştıA little over a week at Quesma, and a lot has already happened. We welcomed Michał Warda, founder of qforge.studio for an inspiring visit. We launched a new website for the product we are already dogfooding internally: a way to understand what coding agents actually do, what they cost, and what outcomes they produce. And we published three technical deep dives: • Piotr Migdał on how quantization hurts factual knowledge nonlinearly • Piotr Grabowski on why retrying interrupted LLM requests introduces length bias • Bartosz Kotrys on how seemingly sensible AI orchestration backfired 4 ways More is just around the corner. Links in the comments.
-
Jacek Migdał bunu beğendiJacek Migdał bunu beğendiIf you’re an LP and have been backing these Berlin (!) emerging managers since 2022/23, you probably have a lot to be happy about: Lucid Capital l (2023). Johann Nordhus Westarp, Srećko Džeko Conduct · GeneralMind · Bayshore AI Robin Capital l (2023). Robin Haak ARX Robotics · Kombo (YC S22) · Almetra Interface Capital l (2022). Niklas Jansen, Christian Reber Lovable · Superscale AI · RobCo Puzzle Ventures l (2022). Gloria Baeuerlein Rillet · Duna · duvo.ai And much more (emerging) is coming out of Berlin (older vintages). Other emerging managers worth watching: Auxxo Female Catalyst Fund. Dr. Gesa Miczaika, Bettine Schmitz Nucleus Capital. Maximilian Schwarz, Dr. Isabella Fandrych Amino Collective. Manuel Grossmann Discovery Ventures. Dr. Jan Deepen, Stefan Jeschonnek Revent. Otto Birnbaum System.One. Maximilian Claussen Tiny Supercomputer Investment Company. Philipp Moehring Fly Ventures. Gabriel Matuschka A whole new generation of funds is being built in one city — on the shoulders of Cherry Ventures, Earlybird, HV Capital, Point Nine, BlueYard Capital, Visionaries Club, and many others who built the Berlin venture ecosystem before us.
-
Jacek Migdał bunu beğendiJacek Migdał bunu beğendiHad an amazing experience attending the very first conference, #Tokenomicon + #FinOps X in Amsterdam! 🇳🇱 Getting to step into the room with the maintainers and contributors behind the framework, alongside practitioners from all over the world, was inspiring. The energy of the community and the open exchange of ideas made a huge impression on me. With packed tracks across the agenda, there were so many insightful talks happening back-to-back. A few key highlights that stood out to me: • AI spend and token efficiency: direct focus on token usage, understanding that token price alone doesn't show your real spend, and why setting up practical guardrails for experimental AI usage is critical right now. • Unit economics and governance: great sessions on building trustworthy metrics for AI infrastructure, auto-optimizing consumption, and bringing commercial and procurement teams into the conversation early before runaway spend happens. Big thanks to the team at FinOps Foundation and Tokenomics Foundation for putting together such a solid event. Grateful for all the new connections and ideas to take home! 🚀 #FinOpsFoundation #Tokenomicon #FinOpsDayX
-
Jacek Migdał bunu beğendiJacek Migdał bunu beğendiWhile reviewing what we have done for a customer in the last 12 months, one thing became clear: so many security tools generate alerts that are 99% noise. 900 alerts per week on average. Less than 9 real, actionable threats on average. Their entire security team is just 3 people. It would have taken them 6 analysts to figure out which 1% of alerts they should really respond to. AirMDR's cost of providing this service + 24x7 monitoring is less than 1/2 an FTE. And the quality of investigation is much better than what's possible in 15 min per alert. We are finally at a point where we can truly cure security teams of Alert Fatigue. If in 2026 you are still struggling with too many alerts, you are most likely using last-generation tech/processes. If your team is still triaging hundreds of alerts to find the handful that matter, here's how AI-powered MDR compares to the traditional model: https://bit.ly/3Tod9ri
-
Jacek Migdał bunu beğendiJacek Migdał bunu beğendiLast week at Agent Conf was super cool. Great to catch up and actually meet other people building AI agents in person, especially Szymon Rybczak Talks were pretty good too. For me the best ones were Piotr Karwatka on Open Mercato and Jakub Kuzimski from Allegro, both about building an AI software factory (worth watching on YouTube) First time for me, but for sure not the last!
-
Jacek Migdał bunu beğendiJacek Migdał bunu beğendiMalgorzata Piotrowska, Cezary Piwowarczyk, and I had a great time at the first Tokenomicon in Amsterdam, organized by the Tokenomics Foundation and FinOps Foundation. We wrote a recap of our favorite talks, practical lessons on AI costs, and where coding-agent billing still needs work. One takeaway: fewer tokens don’t always mean a smaller bill. You can read it here 👇 https://lnkd.in/dUst3Biz It was great to finally meet and talk with Stephen Arthur and Kevin Emamy from the FinOps and Tokenomics Foundations in person. I really enjoyed the conference format, and Henrique Amorim deserves a special mention for the energy he brought to the stage. A real talent! Thanks also to Łukasz Kubicki, Rafael S., ☁️ Jean Latiere ☁️, Ali Rahi, Yuriy Prykhodko, Rense Siegmund, Andrea Blasioli, Jan C. Simons, Paweł Faszyna, Sebastian Amrogowicz and Magdalena Rapacz-Matras for the great conversations. Hope to see you all again soon!Token ecomonics in Amsterdam: Inside the first Tokenomicon - Quesma BlogToken ecomonics in Amsterdam: Inside the first Tokenomicon - Quesma Blog
-
Jacek Migdał bunu beğendiJacek Migdał bunu beğendiFirst of all - wow, Tokenomicon FinOps X Amsterdam was an incredible event. Thanks to J.R. Storment for putting the trust in me to help facilitate the Tokenomics Foundation member working groups and then asking me to talk about that work - on the big stage! That was the product of many weeks of group effort, and I'm so happy that it's all live now on https://lnkd.in/g3MFBZ9A The analysis of the first State of Tokenomics survey has also just been released! I'm thrilled to be able to point to the top findings and challenges you reported having and have the opportunity to talk about solutions. There is a ton more work to do as we dig into the results - together, and in the open. https://lnkd.in/gRjw6d_K It was also so cool to finally meet (some of) my teammates for the first time as coworkers in Amsterdam - Kevin Emamy Steve Trask Jessica Brandes Henrique Amorim Ishita Vyas Dean Oliver James Newton Sam Benson Sara Whitehead Joe Eletto Rob Martin AJ Witt I had a number of deep and insightful conversations with faces old and new, and I love the energy you're all bringing to Tokenomics. My top takeaway is the validation that we're moving in the right direction: actively prioritizing and working on the biggest problems you all face. Let's keep the conversation going - please continue to share questions and ideas with me here or in Slack!
-
Jacek Migdał bunu beğendiJacek Migdał bunu beğendiChainloop v1.112.0 is out, with many good features around agentic security and AI governance - AI Code Analysis preview - Improvements to prompt-to-production tracking, aiming for end-to-end traceability - Finding deduplication strategy, security fixes and more! https://lnkd.in/eHtUuPhnAI Code Security Analysis - Chainloop DocumentationAI Code Security Analysis - Chainloop Documentation
-
Jacek Migdał bunu beğendiThis week at Quesma: Malgorzata Piotrowska, Cezary Piwowarczyk, and Bartosz Kotrys attended the sold-out Tokenomicon + FinOps X in Amsterdam, alongside 500+ practitioners who own AI budgets. Our technical benchmarks are gaining traction. Jacek Migdał spent two days with peer founders at the Inovo.vc Magia Summit. Piotr Migdał taught the next generation of engineers at the Model Trainers workshop. VCs are knocking too. Piotr Bukanski from Acurio Ventures visited our Warsaw office. Meanwhile, we merged many PRs and dogfooded our product. Next month, we will launch our product: analytics for agentic coding and go on a US tour: SF Bay Area, Oct 19 to 30, with a booth at AGNTCon + MCPCon in San Jose on Oct 22 and 23.
Deneyim
Eğitim
-
Uniwersytet Warszawski
-
-
Faaliyetler ve Topluluklar:organizer at Flaszki - series of events, lightning talks (TED alike)
Qualified with a scholarship for best candidates.
Master thesis: "Low-power Wireless Networks for Mobile Applications" (hardware, C) -
-
-
-
-
-
-
-
Diller
-
English
Tam profesyonel yetkinliği
-
Polish
Ana dil veya ikinci dil yetkinliği
-
German
Başlangıç düzeyinde yetkinlik
Alınan tavsiyeler
5 kişi, Jacek Migdał adlı kullanıcıyı tavsiye etti
Görmek için katılınJacek Migdał adlı üyenin tam profilini görüntüleyin
-
Ortak tanıdıklarınızı görün
-
Başka biri aracılığıyla tanış
-
Jacek Migdał ile doğrudan iletişime geçin
Diğer benzer profiller
Diğer gönderileri keşfedin
-
Vincent Vanhoucke
Waymo • 21 B takipçi
You can think of agentic engineering as just another compiler layer, which sits on top of your C++ compiler, which sits on top of LLVM, etc... But the analogy breaks down, because prompts are ephemeral, and don't stick as the literate description (in Don Knuth's sense of "literate programming") of the code. The main reason is that your "compiler," the LLM, is neither hermetic nor deterministic. I do wonder if we're in a wrong local optimum of vibe coding as a result: imagine a world in which you check in your prompts as the source of truth, and the resulting code is what's ephemeral, or a blend of both like in CWEB. This would bring some rigor and sharing capabilities to the exercise of prompting agents, and elevate the entire codebase by one abstraction layer. Another big challenge is that the results of a prompt are non-local, and as a consequence, using them as a linear description of the code is not particularly useful: it's more like a version tree. It is be super interesting to think about what a prompting-aware 'git' would look like, and the role of comments vs. prompts in this environment.
120
14 Yorum -
Reid Pinchback
Specialties: Full-stack… • 1 B takipçi
Front-loading LLMs with clear specifications can be challenging even for relatively narrow-scoped tasks, like a full-featured production-ready bash script. Encouraging an LLM to ask questions if any aspect of the specifications is incomplete or confusing, often as not, gets ignored. Adding asterisk-delimited text sorta-kinda-sometimes nudges it to ask. Adding ampersand characters as delimiters works even better. One small problem: Now it won't *STOP* asking questions. "I noticed you use the number 1. Is that number 1 really a number 1? Please confirm before I proceed." (makes prompt change: "but encouraging you to ask questions does not mean ask questions that the specification obviously already answers in three different places") It's the tedious part of so-called prompt engineering. LLM behavior flips from "ignoring you" to "over-reacting" very easily. You can almost feel it jumping a low-probability transition from one sticky cluster in the weights to another cluster. Arguably I think it's the biggest gap we have in making this be actual engineering instead of what it tends to be in practice: permute your instructions until you luck out. What we actually need is an observability mechanism that identifies the sticky clusters and measures how specific instructions move you between them. I've got some articles coming out soon that uses Markov chains to show how some of the math around this works. LLM inference isn't absolutely the same thing but the underyling mechanics are pretty close.
2
-
Adam Warski
VirtusLab • 9 B takipçi
What kind of guidance does an LLM need to write a direct-style #Scala 3 application? At the baseline - not a lot, e.g. Claude is quite good both in Scala 3 & direct-style. For finer details - some additions to the prompt might be useful. Read more: https://lnkd.in/dxwy-WdD
20
5 Yorum -
Victor Fei
Ormi Labs • 3 B takipçi
Any Subgraph provider claims 99.9% up time is most likely misleading you (although not intentionally). {⏰ 19 days until Alchemy Subgraph shut down ⏰ } I have seen graph nodes take up to two hour to reconcile a re-org. Since this is technically "outside" the provider's control, nor can they actually intervene, they don't count it as downtime. Subgraphs still technically "Operational" although there is a massive delay due to chain re-org. 🪫 They maintain their 99.9% claim, but your data is stale, user transactions fail, and your community on Discord is mad. 🤯 You must clarify this with your new Subgraph provider when moving away from Alchemy Subgraphs Most Blockchain Backend Engineers assume uptime is binary: it is either up or down. The truth is, blockchain availability is never one-size-fits-all, yet Subgraph providers rarely tailor their re-org configurations per chain. Re-org impact is the grey area providers don't mention. That is why your dApp must have a system to respond to these nuances: 1️⃣ Define Freshness Targets Per Chain I see teams treat every chain the same, but block times vary wildly. You must define specific latency targets for each network. e.g. If Arbitrum-one falls behind 20 blocks for 5 minutes, you need to know immediately. 2️⃣ Set Granular Alert Thresholds Per chain Don't rely on generic uptime pings. Your alert thresholds should be set per chain. If Monad stalls, it probably shouldn't trigger the same alert urgency as a delay on Mainnet unless the threshold is crossed. Trigger the on-call person only when it matters. 3️⃣ Implement Alerts for Re-orgs I’ve seen projects fly blind during re-orgs because they assumed the Subgraph would just "catch up" silently. You need an explicit alert for when a re-org is detected. Deep re-orgs can wreck data consistency if you aren't watching. sometimes it is not automatically recoverable you have to inform your subgraph provider. e.g. XLayer in June 2025, from 3M block re-org introduced block 7M to re-org causing all subgraph providers to fail Lastly your new provider will definitely have different re-org threshold settings than Alchemy. This means your recovery time could be longer than you expect, leaving users staring at stale data. -- If you are a Blockchain Backend Engineer looking to migrate your Alchemy Subgraph to a new provider with zero disruption to your dApp and want more educational resource. Get my 📑 5-Step Alchemy Subgraph Migration Checklist 📑. Follow the link in the comments below to grab your complimentary copy 👇👇
4
1 Yorum -
Jeffrey DeCoux
36 B takipçi
Vijay Bendigeri, Ph.D., excellent post. The truth is, this new multi-trillion-dollar Industry requires INTELLIGENT INFRASTRUCTURE to address supply chain logistics, optimization, orchestration, reduced emissions, Vision Zero, and enabling Autonomy 2.0. Exceptional work has advanced the vehicle with sensors, software, computing, AI algorithms, and safety features. It is far more than the vehicle. Exciting times for our nation, Building Our Next Golden Age. The national buildout of Intelligent Infrastructure Economic Zones will revitalize infrastructure in the United States. Supporting NextG for DoD and Civil, FAA NextGen, Intelligent Transportation, Golden Dome, resilient grids, precise positioning, and autonomous systems. ODC AI-RAN is the convergence of AI and the physical world. Autonomy Institute https://lnkd.in/giehikzr
7
-
Umar Lateef
Rectify • 9 B takipçi
A Claude Code tutorial just hit 4.8M views on X. The author? Eyad Khrais - ex-Amazon and Disney dev, now a CTO building enterprise agents. I read his playbook. Here's what actually matters: 𝟭. 𝗧𝗛𝗜𝗡𝗞 𝗕𝗘𝗙𝗢𝗥𝗘 𝗬𝗢𝗨 𝗧𝗬𝗣𝗘 Press Shift+Tab twice to enter plan mode. This 5-minute step saves hours. Don't say "Build an auth system." It gives too much creative freedom. Instead: "Build email/pass auth using the User model, Redis sessions (24h expiry), and middleware for /api/protected." Architecture first. Automation second. 𝟮. 𝗖𝗟𝗔𝗨𝗗𝗘.𝗠𝗗 𝗜𝗦 𝗬𝗢𝗨𝗥 𝗟𝗘𝗩𝗘𝗥𝗔𝗚𝗘 𝗣𝗢𝗜𝗡𝗧 Claude reads this file every time. But attention is scarce: 1. It can follow ~150 instructions max 2. System prompts use ~50 3. Every new line competes for attention Keep it surgical. Explain WHY, not just WHAT. • Bad: "Use TypeScript strict mode" • Good: "Use strict mode because we have bugs from implicit any types" Context enables better judgment calls. 𝟯. 𝗖𝗢𝗡𝗧𝗘𝗫𝗧 𝗗𝗘𝗚𝗥𝗔𝗗𝗘𝗦 𝗔𝗧 𝟯𝟬% Opus 4.5 has a 200K window, but quality drops at 20-40% usage. If Claude compacts and outputs garbage, the model was already degraded. The fix? The copy-paste reset: • Copy key info from terminal • Run /compact • /clear context • Paste back only what matters Fresh context preserves intelligence. 𝟰. 𝗪𝗛𝗘𝗡 𝗖𝗟𝗔𝗨𝗗𝗘 𝗚𝗘𝗧𝗦 𝗦𝗧𝗨𝗖𝗞 If you've explained it 3x, stop. More explaining won't help. Change the approach: • Show examples instead of telling • Break complex requests into pieces • Reframe: "Implement state machine" vs "Handle transitions" Recognize loops early. Don't fight degraded context. 𝟱. 𝗨𝗦𝗘 𝗧𝗛𝗘 𝗥𝗜𝗚𝗛𝗧 𝗠𝗢𝗗𝗘𝗟 • Opus: Planning, reasoning, architecture • Sonnet: Execution, refactoring The winning workflow: Opus plans, Sonnet executes. Your CLAUDE.md ensures they share constraints. 𝟲. 𝗕𝗨𝗜𝗟𝗗 𝗦𝗬𝗦𝗧𝗘𝗠𝗦, 𝗡𝗢𝗧 𝗢𝗡𝗘-𝗦𝗛𝗢𝗧𝗦 Use the -p flag for headless automation. This enables scripting for PR reviews and ticket responses. The flywheel: Claude makes a mistake → Review logs → Update CLAUDE.md → Claude improves. This compounds. Read Eyad's full article here: https://lnkd.in/er8_YTyg
520
21 Yorum -
Lee Gonzales
BetterUp • 4 B takipçi
I ran `apk add python3` inside my browser tab yesterday. Not a simulation. A full Alpine Linux distribution with a real package manager, executing real x86 binaries through JavaScript emulation. This is Apptron, and it broke several assumptions I didn't know I had. The economics are backwards: Traditional cloud IDEs (Gitpod, Codespaces, Replit) run compute on their servers. You pay for it. Apptron inverts this—all the heavy lifting happens in YOUR browser. The backend just handles file sync and networking. Part of the infrastructure runs on Cloudflare's free tier. The cost of providing a full Linux environment to users is essentially zero because users bring their own computers. It's not emulating Linux. It's running Linux: At the core is v86, a JavaScript x86 CPU emulator with JIT compilation. On top sits Wanix, a WebAssembly kernel framework bridging the emulated environment with browser APIs. Real, unmodified 32-bit x86 binaries execute in your browser. Why this matters: → Interactive documentation with actual Linux environments → Programming education without "works on my machine" → Legacy software preservation via URL → AI agent sandboxing (the use case I keep coming back to) Browser as OS substrate. User as infrastructure. I don't know if this project will take off. But the pattern it demonstrates is something I hadn't considered. Now I'm wondering what else we've been overcomplicating. 🔗 Try it: https://apptron.dev 🔗 Source: https://lnkd.in/g2JXtD7q Full writeup on my Substack (link in comments). #WebDevelopment #Linux #CloudComputing #DeveloperTools #OpenSource #WebAssembly #AI #SoftwareEngineering
8
3 Yorum -
Andy Bell
Inertia Labs • 1 B takipçi
Hot take: most "AI agent" systems today are not usable in any trust-sensitive environment. They're non-deterministic, non-reproducible, and non-verifiable. You can't audit them, you can't prove what they did, and you can't safely let them touch anything that matters. And yet we're trying to use them for financial flows, on-chain interactions, and autonomous coordination. If agents are going to operate in adversarial or trust-minimised systems, **"it probably did the right thing" isn't good enough**. So I built something different. luai flips the typical agent loop. Instead of the LLM calling tools in a back-and-forth ReAct loop, the LLM generates a complete Lua program in a single call. That program then runs in a sandboxed, deterministic VM — making tool calls, processing results, and returning structured output — all at zero additional LLM cost. Every tool call and result is captured in a transcript that can be replayed, verified, and optionally zk-proved. Why Lua? Lightweight, embeddable, and sandboxed by nature — no filesystem, no network, no OS access. Tools are the only external interface. The goal is simple: treat agent execution as **computation you can verify**, not behaviour you hope is correct. My guess is that most current agent frameworks won't survive contact with real-world constraints — and we'll end up rebuilding them on top of verifiable execution models anyway. Early days, but the code is here: https://lnkd.in/ejJ9cVuN
10
3 Yorum -
Barzan Mozafari
University of Michigan • 6 B takipçi
Our GenRewrite work is now officially accepted to appear at ACM SIGMOD 2026. Slow query performance is one of the most common complaints in database systems. Part of this is on query optimizers to do better, but a significant portion is caused by poorly written queries - whether from inexperienced users or auto-generated by modern BI tools like Looker, where it's not uncommon for queries to span hundreds of lines. When the gap between what the user wrote and what the optimizer needs is too large, traditional query optimization falls on its face. Whether it's a DBA trying to hand-optimize queries or applying rewrite rules, the problem remains. The traditional approach has been rewrite rules - pattern-matching transformations that replace query fragments with optimized equivalents. But these are rigid. A rule-based rewriter can only optimize patterns it already knows. If your query doesn't match an existing rule, you're out of luck. With LLMs now capable of complex reasoning, the natural question is: can we do better? Can LLMs discover rewrite opportunities that existing rules miss? This is what GenRewrite explores. The challenge is that simply prompting an LLM to "rewrite this query faster" doesn't work well on complex queries. Most outputs either fail to parse or produce incorrect results. GenRewrite introduces Natural Language Rewrite Rules (NLR2s) - human-readable descriptions that guide the LLM toward effective rewrites and enable knowledge transfer across queries. The system learns over time: as it successfully rewrites queries, it accumulates NLR2s and applies them to similar queries based on performance bottleneck analysis. Some examples of NLR2s that GenRewrite discovered: * "Replace the correlated subquery with a precomputed aggregate in a separate CTE" * "Use common table expressions (CTEs) to precompute aggregates" * "Replace a correlated subquery that computes an aggregate with an inline view that pre-aggregates data" We also use counterexample-guided iterative correction to fix semantic and syntactic errors, improving correctness rates significantly - 70% of GenRewrite's rewrites are semantically equivalent to the original, compared to 52% for baseline LLM approaches. On TPC-DS, GenRewrite achieves 2x or higher speedup on 25 queries - 35% more queries than other LLM-based approaches and 160% more queries than LLM-enhanced rule-based systems. Some individual rewrites hit 9x to 24x speedups. On SQLStorm-generated queries that the LLM never saw during training, the gap widens further. Kudos to my PhD student Jie Liu for his great work on this! Full paper: https://lnkd.in/eHBdvw7c #LLM #GenAI #GenRewrite #QueryRewriting #SIGMOD2026
16
1 Yorum