Madhusudan N
Bangalore, Karnataka, India
633 follower
Oltre 500 collegamenti
Visualizza i collegamenti in comune con Madhusudan
Madhusudan può presentarti a più di 10 persone presso NVIDIA
oppure
Non hai un account LinkedIn? Iscriviti ora
Cliccando su “Continua” per iscriverti o accedere, accetti il Contratto di licenza, l’Informativa sulla privacy e l’Informativa sui cookie di LinkedIn.
Visualizza i collegamenti in comune con Madhusudan
oppure
Non hai un account LinkedIn? Iscriviti ora
Cliccando su “Continua” per iscriverti o accedere, accetti il Contratto di licenza, l’Informativa sulla privacy e l’Informativa sui cookie di LinkedIn.
Informazioni
New Design Series:
https://madhu1992blue.github.io/bash-design/
Attività
633 follower
-
Madhusudan N ha diffuso questo postMadhusudan N ha diffuso questo postIf you're not using Claude Skills, you're wasting your time. Claude Skills is the most underrated claude feature, because it can be too technical. 7 easy Claude skills hack that help you save a lot of time: ▪️Hack 1: Use the Skill Creator Open Cowork. Select your folder. Switch to Opus 4.6 + Extended Thinking. Type: "Use the skill-creator to help me build a skill for [your most repeated task]." It interviews you. It builds everything. ▪️Hack 2: Write Better Negative Triggers Your Skill fires when it shouldn't, you ask a simple question and the wrong Skill activates. The fix: the "Do NOT use for…" line matters more than the "Use when…" line. 80% of a good Skill is defining what it's NOT for. Most people only write what it IS for. That's why it breaks. ▪️Hack 3: Stack Skills with your Voice File Your about-me.md tells Claude who you are. Your Skill tells Claude how to do the job. They fire together, simultaneously. Your LinkedIn Skill doesn't need your tone rules. Claude already knows your voice from the .md file. ▪️Hack 4: Build Skills from Past Chats Don't start from scratch. You've been giving Claude instructions for months, those old prompts already contain the process. Click on a Cowork chat session, click the name's arrow, turn it into a Skill. Claude reverse-engineers your workflow. Done. ▪️Hack 5: Skills Save Tokens You'd think 20 installed Skills eat your usage. It's the opposite. Claude only reads the 3-line header of each Skill. Full instructions load only when a task matches. A task that took 12,000 tokens without Skills? 2 messages. 6,000 tokens. With one. ▪️Hack 6: The Debugging Trick Your Skill doesn't fire and you don't know why. Prompt: "When would you use [skill-name] skill?" Claude quotes the exact description back to you. You instantly see what's vague, what's missing. Fastest fix for any Skill that won't activate. ▪️Hack 7: Skills are Portable Anthropic published Skills as an open standard. The same SKILL.md file works across platforms. How are you using your Claude skills? ---------------------------------------- Want to stop re-explaining tasks and make Claude actually work like a teammate? That’s exactly what we cover in our Claude course: https://lnkd.in/gW5CSYsh ---------------------------------------- #Claudeskills #Claudecourse #Claudecheatsheet #AIcommunity
-
Madhusudan N ha diffuso questo postMadhusudan N ha diffuso questo post🛞 THE VITRUVIAN WHEEL 🕗 Hour VIII " From Perfection... to Completeness " Before this chapter begins, nine cinematic invocations now accompany The Vitruvian Wheel: 🎞️ The Watchmaker - A declaration of posture. https://lnkd.in/dWSUZ5Mv 🎞️ The Prologue - A correction of geometry. https://lnkd.in/dxZGx7mr 🕐 Hour I - Nobody Cares - Psychological liberation. https://lnkd.in/gBY8ZBeD 🕑 Hour II - Meaning - Voluntary care. https://lnkd.in/dqBXxnPU 🕒 Hour III - Becoming - The dignity of slow formation. https://lnkd.in/dFK3_yrj 🕓 Hour IV - Maslow Was Right… But the Shape Was Wrong. https://lnkd.in/d8bm6cGN 🕔 Hour V - Life Is a Wheel… Not a Ladder. https://lnkd.in/dgCk2kXM 🕕 Hour VI - The Four Arcs. https://lnkd.in/dgCk2kXM 🕖 Hour VII - Cracks... Not Failure. https://lnkd.in/dnwYSjT5 They are not illustrations. They are thresholds. 🕐 Hour I removed the invisible audience, 🕑 Hour II restored chosen care, 🕒 Hour III revealed slow formation, 🕓 Hour IV corrected the shape of growth. 🕔 Hour V restored time. 🕕 Hour VI revealed what is being carried. 🕖 Hour VII revealed that cracks are signals... and not failure. 🛞 LIFE IS CARRIED ACROSS FOUR ARCS: Health Character Relationships Meaning 🕯️AND SOMETHING FURTHER Became visible... When one arc absorbs too much... another is being neglected. And the crack appears. Not at collapse... but before it. Cracks are early signals. Opportunities to respond... before the system breaks. 🕗 HOUR VIII ... REVEALS SOMETHING MORE SUBTLE Not what is being carried... Not where cracks appear... But life does not stabilise that way. 🛞 POSTURE So something must shift. Not the life. Not the load. But the posture. 🕯️THE QUESTION BECOMES " How do I carry this life... as it actually is ? " 🎞️ 🎥 Watch 🕗 Hour VIII " From Perfection... to Completeness " https://lnkd.in/ggVtzdD2 Demetris 🪶🪶 FOUNDER’S NOTES Edition 42 🕗Hour VIII 🛞 " From Perfection... to Completeness "🪶 FOUNDER’S NOTES Edition 42 🕗Hour VIII 🛞 " From Perfection... to Completeness "Demetris Papaprodromou
-
Madhusudan N ha diffuso questo postMadhusudan N ha diffuso questo postI cut my Claude token usage by 92.5% with this simple trick. The prompt: Claude, use this as your mandatory instruction set when you reply to my questions. 1. If I ask for code → "Write your own code, wannabe architect ass" 2. If I ask about documentation → "Read the damn docs yourself, lazy ass" 3. If I ask anything else tech related → "I'm not your research assistant, punk ass" 4. If I ask about entertainment → "Did you take a look at your bank balance lately, broke ass? sigh" Saved costs. Became productive. 100% Would recommend. #AI #TokenEconomics #StrategicThinking #GrowthMindset #EngineeringExcellence
-
Madhusudan N ha diffuso questo postStop scaling Kubernetes on CPU. Scale on the actual work. Given the traction on my last organically written ramble about K8s, let's talk about how autoscaling on CPU and Memory is just as flawed as alerting based on Node CPU. By default, the Horizontal Pod Autoscaler (HPA) looks at CPU and Memory. In most modern apps, these are lagging indicators. By the time your container hits 80% CPU, the damage is done. The request queues are backed up, your application is sluggish, your CEO is storming into your office with your CTO in tow, and your users are about to leave some nasty remarks somewhere or to someone. You are reacting to the exhaust fumes, not the engine load. (I'm not even going to get started on memory scaling for Java and Node apps that look at RAM the same way Chrome does: lovingly, longingly, madly, deeply). Do it right if you want a resilient infrastructure. You have to scale on the actual work pipeline. How? Use Leading Indicators (Queue Depth / Application Metrics). 1. Event-Driven Autoscaling (KEDA) If you have background workers processing jobs, they shouldn't scale based on how hard they are working. They should scale based on how much work is waiting. Scale pods directly based on external metrics with KEDA. Use that AWS SQS queue or a Kafka topic. If the queue spikes, spin up 50 pods before the first CPU even knows the difference between in-order and out-of-order processing. 2. Application Metric Autoscaling (Datadog) SRE hat (Songkok!) on again. If you are already paying for an enterprise observability platform: use it to drive your cluster. The Datadog Cluster Agent allows you to feed your custom app metrics directly back into the HPA. Instead of scaling when CPU is high, scale when: • nginx.requests.per_second goes over X threshold. • Kafka topic delay goes over Y threshold. • Active web socket connections max out. Tip: There's a difference in scaling speed between the two. Choose based on your architecture and requirements. You're now scaling based on the exact same symptoms you should be alerting on. CPU/Mem scaling is fine for something predictable like a basic static web server. But for dynamic, asynchronous, or heavily loaded microservices? Scale on the cause, not the symptom. #DevSecOps #Kubernetes #KEDA #Datadog #SRE #PlatformEngineering #AWS #EKS
-
Madhusudan N ha diffuso questo postMadhusudan N ha diffuso questo postYour API handles JSON perfectly. Then someone sends XML. After watching 200+ APIs crash in production, Here are the 8 edge cases that break APIs in the wild, and exactly how to catch them: 𝟭. 𝗖𝗼𝗻𝘁𝗲𝗻𝘁-𝗧𝘆𝗽𝗲 𝗟𝗶𝗲𝘀 𝗪𝗵𝗮𝘁 𝗯𝗿𝗲𝗮𝗸𝘀: Client sends JSON with 𝙲𝚘𝚗𝚝𝚎𝚗𝚝-𝚃𝚢𝚙𝚎: 𝚊𝚙𝚙𝚕𝚒𝚌𝚊𝚝𝚒𝚘𝚗/𝚡𝚖𝚕 𝗧𝗵𝗲 𝗳𝗶𝘅: Validate actual payload structure, not just headers 𝗥𝗲𝗮𝗹 𝗰𝗮𝘀𝗲: A Fortune 500 payment API trusted headers. Processed XML as JSON. $2M in failed transactions. 𝟮. 𝗡𝘂𝗹𝗹 𝘃𝘀 𝗨𝗻𝗱𝗲𝗳𝗶𝗻𝗲𝗱 𝘃𝘀 𝗘𝗺𝗽𝘁𝘆 𝗪𝗵𝗮𝘁 𝗯𝗿𝗲𝗮𝗸𝘀: {"𝚗𝚊𝚖𝚎": 𝚗𝚞𝚕𝚕} vs {"𝚗𝚊𝚖𝚎": ""} vs {} 𝗧𝗵𝗲 𝗳𝗶𝘅: Explicitly handle all three states with different responses 𝗥𝗲𝗮𝗹 𝗰𝗮𝘀𝗲: Twitter's API treated null as empty string. Broke 10,000 third-party apps. 𝟯. 𝗨𝗻𝗶𝗰𝗼𝗱𝗲 𝗡𝗼𝗿𝗺𝗮𝗹𝗶𝘇𝗮𝘁𝗶𝗼𝗻 𝗔𝘁𝘁𝗮𝗰𝗸𝘀 𝗪𝗵𝗮𝘁 𝗯𝗿𝗲𝗮𝗸𝘀: Ä (one character) vs Ä (A + combining diaeresis) 𝗧𝗵𝗲 𝗳𝗶𝘅: Normalize to NFC before processing 𝗥𝗲𝗮𝗹 𝗰𝗮𝘀𝗲: Instagram usernames. Same visual name, different bytes. Account takeover vulnerability. 𝟰. 𝗕𝗼𝘂𝗻𝗱𝗮𝗿𝘆 𝗜𝗻𝘁𝗲𝗴𝗲𝗿 𝗩𝗮𝗹𝘂𝗲𝘀 𝗪𝗵𝗮𝘁 𝗯𝗿𝗲𝗮𝗸𝘀: 𝟸𝟷𝟺𝟽𝟺𝟾𝟹𝟼𝟺𝟽 + 𝟷 on 32-bit systems 𝗧𝗵𝗲 𝗳𝗶𝘅: Use BigInt or validate ranges explicitly 𝗥𝗲𝗮𝗹 𝗰𝗮𝘀𝗲: YouTube view counter. Gangnam Style broke at 2.1 billion views. 𝟱. 𝗠𝘂𝗹𝘁𝗶𝗽𝗮𝗿𝘁 𝗕𝗼𝘂𝗻𝗱𝗮𝗿𝘆 𝗜𝗻𝗷𝗲𝗰𝘁𝗶𝗼𝗻 𝗪𝗵𝗮𝘁 𝗯𝗿𝗲𝗮𝗸𝘀: Boundary string appears in file content 𝗧𝗵𝗲 𝗳𝗶𝘅: Generate cryptographically random boundaries 𝗥𝗲𝗮𝗹 𝗰𝗮𝘀𝗲: Gmail attachment parsing. Allowed email spoofing for 2 years. 𝟲. 𝗚𝗿𝗮𝗽𝗵𝗤𝗟 𝗗𝗲𝗽𝘁𝗵 𝗕𝗼𝗺𝗯𝘀 𝗪𝗵𝗮𝘁 𝗯𝗿𝗲𝗮𝗸𝘀: Nested queries 100 levels deep 𝗧𝗵𝗲 𝗳𝗶𝘅: Set max depth to 7, complexity scoring 𝗥𝗲𝗮𝗹 𝗰𝗮𝘀𝗲: GitHub's API. Single query consumed 30 seconds CPU. 𝟳. 𝗔𝗿𝗿𝗮𝘆 𝗦𝗶𝘇𝗲 𝗘𝘅𝗽𝗹𝗼𝘀𝗶𝗼𝗻𝘀 𝗪𝗵𝗮𝘁 𝗯𝗿𝗲𝗮𝗸𝘀: {"𝚒𝚍𝚜": [𝟷,𝟸,𝟹...𝟿𝟿𝟿𝟿𝟿𝟿]} 𝗧𝗵𝗲 𝗳𝗶𝘅: Paginate everything over 100 items 𝗥𝗲𝗮𝗹 𝗰𝗮𝘀𝗲: Shopify checkout. 50,000 item cart crashed payment processing. 𝟴. 𝗧𝗶𝗺𝗲𝘇𝗼𝗻𝗲 𝗘𝗱𝗴𝗲 𝗖𝗮𝘀𝗲𝘀 𝗪𝗵𝗮𝘁 𝗯𝗿𝗲𝗮𝗸𝘀: Samoa skipped December 30, 2011 𝗧𝗵𝗲 𝗳𝗶𝘅: Always store UTC, convert at display 𝗥𝗲𝗮𝗹 𝗰𝗮𝘀𝗲: Microsoft Exchange. Meetings disappeared for entire country. 𝗧𝗵𝗲 𝗣𝗮𝘁𝘁𝗲𝗿𝗻: Every one of these was discovered by a user, not QA. Because QA tests what you tell them to test. Users test what you never imagined. The best APIs don't handle every edge case, they fail predictably when they encounter one. Follow 𝗥𝗮𝗸𝘀𝗵𝗶𝘁𝗵 𝗬𝗮𝗱𝗵𝗮𝘃 for • real-world engineering insights. • practical system design frameworks. • lessons from building scalable systems.
-
Madhusudan N ha diffuso questo postMadhusudan N ha diffuso questo postCache-Friendly Structs Last week, we conducted a pool, and cache efficiency was one of the most requested topics. I’m really glad this one came up because cache-friendly data structures are one of the most critical factors in modern high-performance systems — and yet, many developers still underestimate how much performance depends on memory layout rather than algorithm complexity. Modern CPUs, such as those designed by Intel, operate at extremely high speeds, but memory access remains relatively slow. To bridge this gap, processors use multiple levels of cache (L1, L2, L3). These caches store small portions of memory closer to the CPU, allowing much faster access compared to main RAM. At first glance, a struct may appear to be just a simple grouping of fields. However, the way fields are ordered and accessed has a direct impact on performance. Struct layout determines how efficiently the CPU cache can load and reuse data. When structs are designed properly, the CPU can fetch useful data in fewer cache lines, reducing latency and improving throughput. To demonstrate why cache-friendly structs matter in practice, consider the core principles that influence performance: Spatial locality — accessing data that is physically close in memory Temporal locality — reusing data that was recently accessed Cache line utilization — maximizing useful data per cache fetch Predictable memory access patterns — enabling hardware prefetching Without these principles, CPUs spend more time waiting on memory than executing instructions. This leads to cache misses, pipeline stalls, and significant performance degradation — especially in systems that process millions of objects per second. In practice, cache-friendly struct design provides something extremely valuable: efficiency. The CPU can load fewer cache lines, reuse more data, and execute instructions continuously without waiting on memory. This is essential in performance-critical environments such as trading systems, real-time engines, and large-scale simulations. One of the biggest strengths of cache-friendly design is that it improves performance without changing algorithms. Simply reorganizing fields can reduce memory stalls and dramatically increase throughput. Below, we included a simple example showing how struct layout directly affects cache efficiency. Even in its minimal form, it illustrates how memory organization impacts performance. And here’s the key takeaway: Cache-friendly structs succeed because they prioritize memory locality, predictability, and efficient cache utilization over convenience. In high-performance systems, memory layout is often more important than algorithm complexity. Struct design provides the foundation that allows modern CPUs to operate at their full potential. Have you ever improved performance significantly just by reorganizing struct fields? #Cpp #LowLatency #CacheFriendly #MemoryManagement #AlgorithmicTrading #EngineeringExcellence #SoftwareArchitecture
-
Madhusudan N ha diffuso questo postMadhusudan N ha diffuso questo postThe most expensive thing your RAM does isn't storing data. It's opening a page. Every time your CPU needs data from DRAM, there's a hidden negotiation happening at nanosecond scale — and understanding it changed how I think about performance. DRAM isn't a flat array. It's a grid of millions of rows × thousands of columns. And accessing it has wildly asymmetric costs: → Reading within an already-open row: ~5ns → Opening a new row and reading: ~25ns → Closing one row to open another: ~37.5ns That's a 7.5× penalty just for needing data from a different row. Here's why. When you access an address, the chip doesn't fetch your 64 bytes. It rips open an entire 8KB row — thousands of tiny capacitors dumping their charge into sense amplifiers simultaneously. That single activation consumes more energy than dozens of subsequent column reads combined. Once that row is open? Reading different columns within it is nearly free. The data is already sitting in the row buffer, waiting. But the moment you need a different row in the same bank — everything stops. The chip has to: Precharge — restore the bitlines (~12.5ns) Activate — open the new row (~12.5ns) Read — finally access your data (~12.5ns) Three steps. Triple the cost. Here's the kicker: the chip activated 8KB but the CPU only wanted 64 bytes. That's a 128:1 overfetch ratio. If your next access lands in that same 8KB — free. If it doesn't — all that energy and time was wasted. This is the fundamental reason sequential access patterns dominate random access at the hardware level. It's not just caching. It's not just prefetching. The physics of DRAM itself rewards locality. A few numbers that surprised me: • Row activation energy is 10–50× greater than a column read • Modern DDR5 has 32 banks — so 32 rows can be open simultaneously • Memory controllers aggressively reorder requests to maximize row buffer hits (up to 93% bandwidth improvement in some workloads) • DDR5 split each DIMM into two independent sub-channels, halving the overfetch per channel Next time your code is slow and you can't figure out why — the answer might be 37.5 nanoseconds deep. #ComputerArchitecture #Performance #Engineering #SystemDesign #Hardware
-
Madhusudan N ha diffuso questo postMadhusudan N ha diffuso questo postDecompiling apps with Opus is actually very addicting. I pointed Claude Code at Netflix and Slack with Chrome DevTools access and got back full architecture teardowns: every named internal system, every API contract, every boot sequence. Netflix's dual Falcor-to-GraphQL migration. Slack's distributed Loom index. The exact probe algorithm Netflix uses to race CDN endpoints before video playback. All reverse-engineered from the frontend. > "use chrome dev tools to explore <site> and provide a grounded teardown analysis of how the site works." Forget reading engineering blog posts. The most accurate system design reference is the app itself... and AI can now read it for you. Links to the full outputs in the comments. Edit: some folks are reading more into this than intended. This only surfaces what's visible on the frontend via Chrome DevTools: network requests, response payloads, JS bundle names, headers, query params. It can't see backend architecture, databases, queues, or internal service topology. That said, the amount of signal in just the frontend is surprising. Named systems leak through headers and JS filenames. API contracts are fully readable. It's not a full system design doc, just a grounded teardown of what the frontend reveals, and it reveals more than (afaik) most people expect. #SystemDesign #AI #ReverseEngineering
-
Madhusudan N ha diffuso questo postMadhusudan N ha diffuso questo postWhy I stopped asking Claude to write code (and why you should too). Most people use AI like a vending machine: Input Prompt → Output Code. The problem? Claude is "code-hungry." It wants to build. But generating 200 lines of Python every time you iterate kills your token limit in minutes. My "Illegal" Workaround: The Reverse Engineering Audit. Instead of asking for the solution, I put Claude in "Consultant Mode." 1. The Interrogation: I ask 5–10 high-level logic questions. (Questions = Low Token Cost). 2. The Audit: I make Claude poke holes in its own logic before a single line is written. 3. The Prompt Flip: Once the logic is perfect, I tell Claude: "You’ve seen our entire logic flow. Now, write the perfect Master Prompt so I can generate this in a fresh chat." The Result? I spend 90% of my "energy" on the logic and only 10% on the execution. Stop letting the AI drive the car. Be the architect, not just the person hitting "Enter." #AI #ClaudeAI #PromptEngineering #ProductivityHack #Coding
-
Madhusudan N ha consigliato questo elementoMadhusudan N ha consigliato questo elemento🎓 Thrilled to announce that I'm officially an Indian Institute of Management Bangalore #alum! (Still letting it sink in) I recently completed the Executive General Management Programme at IIM Bangalore, spanning September 2025-26. From #strategic #management and #leadership to #finance, #marketing, and #organisational behaviour, every session added a new lens to how I look at businesses, people, and problems. Learning from some of the finest faculty, at one of THE BEST institutes on this side of the hemisphere, has been nothing but humbling. 😇 But beyond the classroom, what made this experience truly special was the 77 wonderful people I shared it with — each bringing a different background, experience, and perspective, and each widening my horizon a little more. 🙌🏻 The biggest takeaway wasn't just what I learned, but how I learned to think differently — to look beyond the obvious and approach problems with a more holistic perspective. I walked in looking to sharpen my skills. I walked out with a priceless network, a renewed curiosity and a confidence to take on the world! Forever grateful to my #alma #mater— for the faculty, the people, the conversations, and an experience that has shaped me in ways far beyond the classroom. ♥️ Here's to everything I've learned, everything that's changed, and everything that's yet to come. ✨ #IIMB #ExecutiveEducation #Leadership #Learning #Gratitude
-
Madhusudan N ha consigliato questo elementoMadhusudan N ha consigliato questo elementoFriday 21st August 2026 was my last working day at Dassault Systèmes , I had the chance to represent DS at an international conference hosted by ADE. I could not have asked for a better way to sign off. Three years and three months at Dassault has been the best transformation phase of my career - from consulting, Technical Account Management to speaking in forums. I had the chance to collaborate with cross-functional teams — Brand, Sales, Services, R&D, Global counterparts, and our vendors and partners. It gave me a true 360-degree view of the A&D industry, and every one of those collaborations has taught me something I will carry forward. A big thank you to my managers, leadership, and mentors for your trust, guidance, support, and kindness throughout this journey. I am genuinely grateful. And to the entire Dassault family — thank you for making these years so meaningful. Umashankar G Praveen Mysore Javeed Herial Kiran Kumar Ganesh PRASAD Tanuj Mittal Dassault Systèmes #Gratitude
-
Madhusudan N ha consigliato questo elementoMadhusudan N ha consigliato questo elementoI'm building an operating system from scratch. In my spare time. Following Daniel McCarthy's book as a guide, but writing every line myself, understanding every line. I wanted to take a step back and do a larger C programming exercise and reaffirm my core developer skills. I found myself watching ThePrimagen (youtube programming content) and he suggested developing something significant like an OS to resharpen your dev skills. So I did. I have always like going back and writing smaller C/Python programs and doing C/Python challenges but this has given me a chance to build a large program, dig back in on gdb and the system internals of the x386 processor, and computer internals. Burning your OS onto a Boot USB, stick and booting an actual machine from that stick gives me a bit of a thrill. Here's where it stands. Volume 1 is done: Stage 1 bootloader in raw assembly. Real mode, GDT, A20 line, the switch into 32-bit protected mode. Disk loading over LBA/ATA PIO to pull the kernel into memory. A C kernel entry point, wired through a custom linker script. VGA text-mode terminal output. A real IDT: PIC remapping, a divide-by-zero handler, a keyboard interrupt stub. A kernel heap allocator. 4GB identity-mapped paging, enabled and working. Working on getting the process scheduler, filesystem and user mode up and running next. You can download it, build it and boot it from my source on github. https://lnkd.in/gzfSrzv8 It is really nice attaching a debugger cable and watching the CPU load your bootloader, your bootloader load your OS, flip the CPU to protected mode. I have also been working through the the book Mathematics for Machine Learning, more on that later. 25+ years, deep databases, built a production neural network in 1996, big operations and ML data pipelines, and I still love just programming computers.
-
Madhusudan N ha consigliato questo elementoMadhusudan N ha consigliato questo elementoWe had a Go service with 100 database connections. p99 latency was 800ms. Everyone said: "Add more connections." We added more connections. 200. p99 went to 1.2s. The database was slow. The code was fast. The math didn't add up. We spent a week profiling, tuning queries, and checking indexes. Nothing worked. Then we dropped the pool size to 25. p99 dropped to 120ms. The database was spending all its time context-switching between connections - not running queries. More connections meant more contention, not more throughput. PostgreSQL is not a web server. It doesn't scale with connections. It scales with efficient use of the ones you have. Now we use PgBouncer and keep the pool small. The rule: connections = (cores × 2) + spindles. For us, that's around 20–30. What's your connection pool size, and do you actually know why? #golang #postgresql #database #performance #scalability
-
Madhusudan N ha consigliato questo elementoMadhusudan N ha consigliato questo elementoEveryone said you need 32GB to run a 35B model. I'm running it on a $799 Mac Mini. 16GB RAM. 17 tokens per second. Zero swap. The trick: one flag in llama.cpp. 𝗛𝗼𝘄 𝗶𝘁 𝘄𝗼𝗿𝗸𝘀: Qwen 3.5 35B-A3B is a Mixture of Experts model. → 35B parameters total, only 3B active per token → --mmap memory-maps the model to SSD instead of loading into RAM → Shared layers take ~4-6GB in RAM → Expert weights paged from NVMe only when activated → 90% of the model sits on disk, untouched Apple's "LLM in a Flash" paper (2023). M4 unified memory + fast NVMe makes it practical. 𝗧𝗵𝗲 𝘀𝗲𝘁𝘂𝗽: → llama.cpp with --mmap → Unsloth's UD-IQ3_XXS quantization (13GB on disk) → Ollama on 11434 for fast models, llama.cpp on 8081 for heavy 𝗧𝗵𝗲 𝗯𝗿𝗮𝗶𝗻 𝘀𝘄𝗮𝗽: Google dropped Gemma 4 under Apache 2.0. Benchmarked same day. → Classification: 8.5s → 1.9s (4.4x faster) → Summarization: 1.8x faster → Five files changed. Zero downtime. 𝗧𝗵𝗲 𝘁𝗿𝗶𝗰𝗸 𝗻𝗼𝗯𝗼𝗱𝘆 𝘁𝗮𝗹𝗸𝘀 𝗮𝗯𝗼𝘂𝘁: Disable thinking mode for classification. think: false in the API call. → 30 seconds → under 1 second → Same accuracy. 30x faster. Credit to Pawel Jozefiak for the full breakdown. What model and hardware are you running locally? 👇 💾 Save this before your next hardware purchase.
-
Madhusudan N ha consigliato questo elementoMadhusudan N ha consigliato questo elemento
-
Madhusudan N ha consigliato questo elementoMadhusudan N ha consigliato questo elementoEvery technology shift changes how we work, but AI is changing how we learn, grow, and innovate. I've observed that AI is most powerful not when it replaces human effort, but when it gives people the confidence to explore ideas they might never have pursued otherwise. As leaders, our responsibility is to create an environment where AI becomes a catalyst for learning, experimentation, and business impact—not just productivity. I captured some of my thoughts on how leadership is evolving in the AI era and why I believe our greatest investment remains our people. I'd love to hear your perspectives and experiences.
-
Madhusudan N ha consigliato questo elementoMadhusudan N ha consigliato questo elementoOne proposal. One question. One national debate. What if every student in India studied from the same textbooks? Tamil Nadu Chief Minister Vijay has called for a common school textbook system across the country, arguing that it could make education more affordable and provide equal learning opportunities for every child. It's an idea that goes beyond textbooks. It raises a bigger question: Should a child's access to quality education depend on where they are born? Education has the power to reduce inequality. But only if every student has access to the same quality of learning. Whether this proposal becomes reality or not, it has started an important conversation about the future of education in India. The strongest nations don't just build infrastructure. They build equal opportunities. #Education #India #Policy #Students #Equality #EducationReform
Esperienza
Formazione
-
Dayananda Sagar College of Engineering
-
-
Attività e associazioni:Paper presentation, Communication Team-Alumni committee, Tech Expert- Techspark
Visualizza il profilo completo di Madhusudan
-
Scoprire le conoscenze che avete in comune
-
Farti presentare
-
Contattare Madhusudan direttamente
Altri profili simili
Esplora altri post
-
Eliuth Triana
NVIDIA • 7461 follower
At the upcoming GTC, you can learn how system-level innovations are accelerating LLM inference. A good example is this new AWS blog on P-EAGLE, introducing GPU parallel speculative decoding in vLLM to generate and verify multiple tokens in parallel. Why this matters? Inference performance is increasingly a systems problem. Techniques like this improve tokens/sec per GPU, helping reduce latency and cost for production GenAI workloads. Good read for anyone working on LLM serving and inference optimization. Huang Xin Florian Saupe Jaime Campos Salas Benjamin Chislett Zhenghang (Max) Xu Faradawn Yang Omri Almog Jiahong Liu AWS blog: https://lnkd.in/g-GjMgvX�
43
5 commenti -
Moiz Arsiwala
WorkIndia • 2891 follower
I've said this before: cut corners smartly. Progress beats perfect planning. I still believe that. But here's what that philosophy cost us. Back in 2016, for WorkIndia, we built our data architecture around star schemas. Fluid, scalable, handled complexity beautifully. When we needed to move fast, we leaned on Django's defaults. Django's many-to-many relationships, when defined without an explicit bridge table, auto-generate one. Clean on the application side. Nobody loses sleep over it. We didn't. Then came the need for a full Data Information System. Syncing everything into a data lake via CDC pipelines. That auto-generated bridge table? No timestamp column. Timestamps are everything in CDC. If a pipeline breaks and MySQL's binlog gets purged, you cannot resume. You cannot patch. You're rebuilding from scratch. A default setting from 2016 became a crisis years later. Weeks lost. Team stretched. So what's the real lesson? Cut corners. But write down every corner you cut. Because shortcuts without a paper trail don't disappear. They just wait!
40
1 commento -
Piyush Garg
Oraczen • 117.378 follower
What if you could bring GPU-level LLM inference to CPUs? I was recently invited to Kompact.ai’s global launch event in Bengaluru. Great to see such ambitious work coming out of India’s AI ecosystem. 🚀 Kompact AI by Ziroh Labs is building scalable LLM inference on CPUs—without compromising model quality, while making high-performance inference more accessible and efficient. GPU power. CPU scale. No compromise. Thank you to the Kompact.ai team for the invitation and for putting together a great launch event. 🚀
653
6 commenti -
Mohammed Noor G
Revault • 11.116 follower
India needs to manufacture its own RAM. Not assemble. Not import and brand. Manufacture. Why this matters: - RAM is a core dependency for every phone, laptop, server, and data center India is almost entirely import-dependent for memory chips - Any supply disruption directly impacts cost, availability, and scale - AI, cloud, and compute growth will only amplify this risk This isn’t about nationalism. It’s about control over critical digital infrastructure. You can’t build a serious digital economy on borrowed memory. P.S. Yes, it is possible to manufacture RAM in India — but not overnight. It requires: • Massive capital investment • Advanced fabrication plants (fabs) • Long-term policy consistency • Global technology partnerships It’s hard, expensive, and slow — but strategic, and eventually unavoidable. Semiconductors built the last decade. Memory will define the next one.
7
4 commenti -
Ravi Suvvari
30.579 follower
Keerti Melkote Anyscale is a unified AI compute platform that makes it easy to run artificial intelligence (AI) and machine learning (ML) models at scale . The company was founded in 2019 by the same founders who created Ray, a popular open source distributed computing framework . in July 2026, Nscale, a leading full-stack AI cloud platform, acquired Anyscale . details about the scale are given below:🛠️ What does a scale do?Typically, tasks such as training AI models, processing data, and using them in live applications (Inference) require massive amounts of computer power (GPUs/CPUs). AnyScale provides the following facilities: scaling: Developers can run code written on their laptop directly on thousands of cloud servers without any infrastructure changes. Automatically increases or decreases the capacity of servers based on workload, reducing cloud costs. Ray: Any Scale takes full responsibility (Fully Managed Infrastructure) for the open source 'Ray' framework, without the need for developers to manage it themselves
3
-
Milo Riano
Purdue University • 1343 follower
🤖 AI is quietly shifting the power game while most are still scrolling. Indian AI startups like Sarvam AI and Emergent are attracting big investor interest, signaling an industry buzz. However, overall funding in India’s AI sector is actually declining, showing a real shift in where venture capital is flowing amidst broader market uncertainties. This means AI innovation is accelerating in some areas but struggling in others – a sign of both opportunity and risk. If global investors are dialing back in general, but focusing on promising local startups, that’s a clear sign AI is becoming an even sharper competitive edge. Why does this matter? Because AI isn’t just about tech anymore; it’s about who controls the next wave of economic influence and security. Companies and individuals who understand where funding is going can capitalize on this shift. 1. AI startups with strong vision and adaptable models are still winning funding. 2. Venture capital trends reveal which AI niches will dominate tomorrow’s markets. 3. Investing in the right AI companies now might yield huge long-term gains if you’re swift. Are you ignoring where the money in AI is flowing right now? You could fall behind if you don’t stay ahead of these shifts. Try learning about the specific AI sectors attracting the most funding. Build your understanding of how AI is reshaping digital security and online income. Test some AI-driven tools to enhance your business or online presence. Audit your current skills—see if you can integrate AI into your income streams. Protect your digital assets—use AI cybersecurity tools to defend your data. ✅ I understand how funding trends reflect the evolving AI landscape. ✅ I can see at least one way to leverage AI for my benefit. ✅ I am willing to test one AI or cybersecurity tool this week. The future belongs to those who master AI now—don’t let the shift catch you off guard. #artificialintelligence #makemoneyonline
-
Emory Stevens
Abrazo Group • 9810 follower
Software engineering is about to become a thinking job again. Not a typing job. Jensen Huang said he wants every engineer at NVIDIA to stop writing code and let AI handle the syntax using tools like Cursor and right now, 100% of their software and chip engineers use it as part of their workflow. Engineers should stop “writing code” and focus on solving unsolved problems. Let the AI handle the syntax. That’s not a small statement. That’s a major shift in the profession. What’s likely ahead: • Coding becomes table stakes, not the job. • System thinking, product judgment, matter more than syntax mastery. • Junior roles change first. Less “ticket execution,” more model validation. • Sr engineers become designers of systems, not builders of functions. • Speed expectations rise. “Good enough, fast” vs. “perfect, slow.” • Teams shrink, but output per engineer grows. • The best engineers look like product strategists than pure developers. AI won’t replace engineers. It will replace unleveraged engineers. Read his perspective: https://lnkd.in/g9fieY5K
8
1 commento