Xiaoning Wang
Singapore, Singapore
2K followers
500+ connections
View mutual connections with Xiaoning
Xiaoning can introduce you to 10+ people at Amazon Web Services (AWS)
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Xiaoning
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
my proficiency in Python and deep learning algorithms has been instrumental in advancing…
Activity
2K followers
-
Xiaoning Wang reposted thisFlashing back ⏪ to our very first #AWSGenAIinAction story with #SplashMusic! Our team published a detailed technical blog on the solution today: https://lnkd.in/epnf9PGv Here is #SplashMusic CTO Randeep S. Bhatia talking about how they are innovating with generative AI on AWS and how they achieved an 8.6% increase in training throughput while reducing costs by over 55%. https://lnkd.in/e_3fBKWn Sheldon LIU Atanu Roy Nieves Garcia Diez Rossana B. Mahsa Paknezhad Xiaoning Wang Tianyu Liu Xuefeng Liu Daniel Wirjo _________ While the success stories haven't stopped, we're taking a short break ⏸️ from #AWSGenAIinAction stories as we get ready for 𝗔𝗪𝗦 𝗿𝗲:𝗜𝗻𝘃𝗲𝗻𝘁 𝟮𝟬𝟮𝟱. Next few weeks will be action packed, culminating into our biggest, boldest, most comprehensive event of the year, where the global cloud community unites to build the future. I'm excited to see AWS Generative AI Innovation Center (#GenAIIC) builders, customers, and partners on stage showing off their best #AWSGenAIinAction stories! See you in Las Vegas! https://lnkd.in/b_AaCP4Xiaoning Wang reposted this🧵 Real Stories of Generative AI in Action (Feature 1 of a multi-part series, you can access full series at #AWSGenAIinAction) 🎶 Curious how GenAI is impacting music generation? AWS Generative AI Innovation Center (#GenAIIC) recently partnered with Splash Music, an Australia-based company that empowers anyone to create music along with their favorite artists, using AI-powered tools on Roblox and other platforms. Our collaboration focused on the creation of "wemixes", a feature that converts hummed tunes into full musical pieces in any artist's style. 🛑 The Challenge: To enable users to create "wemixes", Splash needed to enhance their infrastructure to efficiently handle petabyte-scale music datasets and LLMs, and to do so cost effectively without compromising their unique approach to model development and music production. 💡 The Solution: Splash Music and AWS collaborated to develop a modern distributed training pipeline, built on Amazon SageMaker HyperPod and leveraging AWS Trainium chips. Specifically, the team utilized a parallelization strategy to enable efficient training of larger models by adding more nodes, with built in support for extended input sequences, as well as NKI Flash Attention and ZeRO-1 optimization techniques to reduce memory footprint and accelerate training convergence. All of this was built using SageMaker HyperPod's SLURM integration, automating multi-node coordination and resource allocation, thus simplifying training management and enhancing operational efficiencies. And of course, no solution would be complete without built in monitoring capabilities to track training progress, resource utilization, and performance metrics. ⁉️ So what? By helping simplify resource management and supporting larger model training, Splash Music opened new possibilities for music creation and innovation. Their new infrastructure helped manage petabyte-scale datasets more efficiently, reducing costs by over 50% while achieving an 8.3% throughput improvement with just 4 Trainium nodes. And for the everyday artists👩🎤👨🎤 ? Over 10K wemixes have been created and 400M+ impressions in five months of beta launch (without everyone having to learn to play🥁🪕🎻🪇🪈🎹🎷, just using their voice🎤🎧)! This collaboration showcases how GenAI-enabled foundational infrastructure for scaling LLMs can help creative teams focus on innovation while maintaining operational efficiency. If you want to see the solution in action: https://lnkd.in/e5z6QTdK #AWS #MachineLearning #ArtificialIntelligence #AIinMusic #GenerativeAI #CloudComputing #SageMaker #Trainium #GenerativeAI #MusicTech #Innovation #CloudArchitecture #TechInnovation #AWSGenAIinAction
-
Xiaoning Wang liked thisXiaoning Wang liked thisI'm thrilled to share that I've joined Databricks as a Senior Director in the AI Forward Deployed Engineering team! This innovative team is helping organizations unlock the full potential of their data and AI strategies, and I couldn't be more excited to contribute to that mission. The team is also growing fast! I'll be sharing some opportunities in the coming weeks including a Senior Manager position. If you're passionate about AI, love working directly with customers to drive transformational outcomes, and want to be part of something special, stay tuned. #Databricks #AI #Hiring #NewRole #DataAndAI #FieldEngineering #FDE
-
Xiaoning Wang liked thisXiaoning Wang liked thisI'm looking for an exceptional leader to join our team. The Generative AI Innovation Center at AWS is hiring a Head of AI Transformation for APJC (Singapore-based). This role leads a team of ML engineers and scientists who work shoulder-to-shoulder with enterprise customers to turn ambitious AI visions into production reality. What makes this role unique? You're not just advising — you're operating. Think GM, not consultant. You'll engage C-suite executives, scope agentic AI architectures, land transformation programs in focused sprints, and expand them across business units. If you're passionate about enterprise AI, thrive at the intersection of technology and business, and want to shape how some of the world's largest organizations adopt AI — let's talk. 🔗 Apply at https://lnkd.in/eM9A27RW #AWS #GenAI #Hiring #AILeadership #APJC #SingaporeHead of AI Transformation - APJC, Generative AI Innovation CenterHead of AI Transformation - APJC, Generative AI Innovation Center
-
Xiaoning Wang liked thisXiaoning Wang liked thisI left AWS two weeks ago (for real this time!). I'm filled with a lot of emotions about stepping away from such a special place, but the main one is deep gratitude. The colleagues, leaders, and customers I've had the privilege to work with over the last 8+ years have taught me so many lessons. About leadership, scale, moving quickly, and making big bets. I am especially proud of the Custom Model & Optimization team, and grateful for the chance to build it over the last three years. We went from an idea at the end of 2023 to a 50+ person global team of some of the finest scientists, engineers, and strategists I have ever worked alongside. We listened to customers and heard loud and clear their desire for optimization in AI, from efficiency to higher accuracy and relevance to more control over their models and systems. I am excited for the next adventure, and forever grateful for my time with AWS. Amazon has been a dream job from my first day to my last, and I could never have predicted all of the amazing things I would get to work on. Thank you, thank you, thank you for the wild, wonderful ride. And thank you Sri Elaprolu, Taimur Rashid, and Francessca Vasquez for your vision and leadership. To the CMO team around the world, you made the last three years the best of my career and I know you are all only just fueling up the rocket ship 🫰🫰🫰
-
Xiaoning Wang liked thisXiaoning Wang liked thisAspire Founders night. A room full of founders, operators and investors. Always good to step out of the day job and hear what people are actually building in Singapore right now.
-
Xiaoning Wang liked thisXiaoning Wang liked thisWhat a day. At the Amazon Web Services (AWS) DC Summit, we announced a $1 billion investment in AWS Forward Deployed Engineering. I could not be more excited about what this means for our customers and partners. The keynote highlighted clearly: the last mile is not a skills problem. It is an acceleration problem. That is exactly what we have seen working alongside enterprise customers for the past three years. They have the talent. They have the ambition. What they need is production engineers embedded alongside their teams who know how to move AI agents from pilot to production in days, not months . That is what Amazon Web Services (AWS) Forward Deployed Engineering delivers. Our engineers embed inside customer teams to co-build production agentic systems with outcome commitments. Every engagement builds a semantic layer specific to the customer, codifying architecture decisions, domain rules, and operational patterns in a governed knowledge graph. The second workflow ships faster than the first. The third, faster still. Customers like Allen Institute, Cox Automotive, the National Basketball Association (NBA), the National Football League (NFL), and Ricoh are already building with us this way. They are not waiting. They are compounding. This is how we close the gap between AI ambition and production reality. Read the Blog: https://lnkd.in/eubdv4md Swami Sivasubramanian Claudio Guglielmelli Rachael Wang Taimur Rashid Matt Garman Shaown Nandi Matt Wood David LevyAWS invests $1 billion to embed AI forward deployed engineers with customersAWS invests $1 billion to embed AI forward deployed engineers with customers
-
Xiaoning Wang liked thisXiaoning Wang liked thisI've spent the day at AWS Summit New York in back-to-back conversations with customers, partners, and analysts. One theme dominated every single meeting: The bottleneck to production agents isn't the model. It's not the prompt. It's not even the data. 𝗜𝘁'𝘀 𝗰𝗼𝗻𝘁𝗲𝘅𝘁. Every executive I sat with today described the same gap: their agents work in isolation. They can't reason across enterprise knowledge — structured, unstructured, real-time — so they're expensive chatbots instead of autonomous systems. Three announcements today that directly address what I heard in every room: 1/ Managed Knowledge Base on Bedrock — native connectors, Smart Parsing across formats, and an Agentic Retriever for complex multi-step queries. Not another RAG wrapper. The retrieval pipeline you don't have to build or maintain. 2/ S3 Annotations — attach up to 1 GB of rich, queryable context directly to objects. Your agent can now query the meaning of a dataset — its lineage, constraints, relationships — before deciding how to act on it. 3/ AWS Context — a self-learning knowledge graph that infers relationships across your data, systems, and business rules. Every agent in your org can tap it. It gets smarter as agents use it. Here's the pattern: most organizations build agents that can answer questions. Fewer build agents that can execute workflows. Almost none build agents that can reason across the entire organization's knowledge autonomously. The gap between "answer a question" and "run a process unsupervised" is almost entirely a context problem. Today's launches close that gap. What's the one context source your agents still can't reach — and what would change if they could?
-
Xiaoning Wang liked thisXiaoning Wang liked this🚀 Super proud to share that our new blog post, co-authored with the Synthesia Research Engineering team, has finally been published. This blog is based on work done as part of an engagement between Synthesia and the AWS Generative AI Innovation Center a few months back. In this blog, we focus on the last step of a video diffusion pipeline, where a large latent video is decoded back to pixel space using the decoder of a 3D video variational autoencoder (VAE). The latent video is decoded one manageable chunk at a time, with the side effect of introducing device-to-host transfers, and therefore GPU stalls, between each decoding step. To fix this and keep the GPU continuously busy, we overlap data transfer with compute using a separate CUDA stream and double buffering. On our example using the HuggingFace Diffusers implementation of the Wan 2.2 VAE running on an Amazon EC2 g7e.2xlarge instance (NVIDIA L40S GPU), we see GPU kernel utilization go from 82% to 99.9% and throughput increase by 8.9%, with a corresponding increase in cost efficiency. Notably, this performance increase was achieved using native precision weights, hence not affecting model quality. The code is available in a public repository so you can reproduce the results and apply this technique to your own workloads. Huge kudos to my co-authors Hanno Bever, Emanuele Levi, PhD, and Moisés Hernández Fernández 🙏 👉 Read the full technical deep-dive: https://lnkd.in/eP_U4Z7S Rohit Thekkanal, Hannah Marlowe, PhDHow Synthesia optimizes generative AI video inference on Amazon EC2 G7e instances | Amazon Web ServicesHow Synthesia optimizes generative AI video inference on Amazon EC2 G7e instances | Amazon Web Services
-
Xiaoning Wang liked thisXiaoning Wang liked thisI’m at #CVPR2026! Just wrapped up the first day of workshops, and it brought back many memories of the computer vision and multimodal lectures I took at CMU. It’s been inspiring to hear brilliant ideas and learn about the latest advances across Efficient Computer Vision, Diffusion Acceleration, Spatial Intelligence and World Model. Excited to spend the rest of the week learning, connecting, and exchanging ideas. If you work also on vision understanding & generation, inference acceleration, or ML system optimization, I’d love to connect!
-
Xiaoning Wang liked thisSpent the last week in Cupertino going deep on the latest from the AWS Neuron SDK lead by Emily Webber and Jim Burtoft with a great group of colleagues — and I'm genuinely excited about where things are heading for Trainium2 & 3. Two highlights: PyTorch Native for Trainium — The team is working on native PyTorch backend integration for Trainium2. Think one-line device change, eager mode for debugging, torch.compile for performance, and standard distributed APIs (FSDP, TP, DDP). Early access is looking really promising. NKI (Neuron Kernel Interface) — When you need peak performance, NKI lets you write custom kernels that target the hardware directly. We went from profiler bottleneck → custom kernel → verified speedup in hours, not weeks — especially with AI-assisted development tooling. If you're running AI workloads at scale and want to hear what's coming for Trainium2, reach out. Cheryl Abundo Wali Akbari Ashwini K. Suji Lee Wayne Toh Nikola Cuca #AWS #Trainium #PyTorch #NKI #MachineLearning #Neuron
Experience
Education
Licenses & Certifications
View Xiaoning’s full profile
-
See who you know in common
-
Get introduced
-
Contact Xiaoning directly
Other similar profiles
Explore more posts
-
Quantum Zeitgeist
19K followers
Graphmend Fixes PyTorch 2 Graph Breaks, Enabling Compilation of 75% More Models and Reducing Fallbacks by 8% GraphMend, a new compiler, automatically eliminates interruptions in PyTorch 2 programs by transforming source code, achieving up to 75% reductions in processing time and boosting performance across several machine learning models. #quantum #quantumcomputing #technology https://lnkd.in/e5kSm8vt
1
-
Aakash Tembhare
Rakuten Symphony • 6K followers
What if the next optimization for LLMs isn’t about computing faster, but looking smarter? I recently explored Declarative Attention an approach where an LLM can decide which part of its context it actually needs to attend to, instead of reading the entire KV cache. The idea is simple: LLM decides where to look → Runtime fetches that context → Model reasons over it. This could be particularly interesting for long-context and agentic workloads. I wrote a short breakdown of the idea and its trade-offs: 🔗 https://lnkd.in/dS8HF3xU Can LLMs learn to decide where they should look before deciding what to think? #AI #LLM #GenAI #Inference #MachineLearning
7
2 Comments -
Pascal Biese
PwC • 86K followers
Agent Memory - is it a critical bottleneck? A new survey from researchers across NUS, Renmin, Fudan, and other institutions offers the most comprehensive taxonomy of agent memory to date. The core insight: existing frameworks like "short-term vs. long-term memory" are insufficient. The field has become fragmented, with researchers using the same terms to describe fundamentally different systems. The authors propose a unified lens of "forms, functions, and dynamics." Memory forms include token-level (explicit, editable text), parametric (encoded in model weights), and latent (hidden states and KV caches). Functions are divided into factual memory (what the agent knows), experiential memory (how it improves), and working memory (what it's thinking about now). Dynamics covers how memories are formed, evolve, and retrieved over time. They identify the following convergence: reinforcement learning is increasingly taking over memory management, moving from hand-crafted rules toward fully learnable systems. The paper argues we may see agents that autonomously design their own memory architectures through RL optimization. All in all, this survey provides a conceptual foundation for building the next generation of agents that can genuinely learn and adapt over time. ↓ 𝐖𝐚𝐧𝐭 𝐭𝐨 𝐤𝐞𝐞𝐩 𝐮𝐩? Join my newsletter with 50k+ readers and be the first to learn about the latest AI research: llmwatch.com 💡
138
9 Comments -
TheNextGenTechInsider.com
1K followers
🌟 New Blog Just Published! 🌟 📌 Optimizing Latency in Private Inference with Coordinate Descent 🚀 ✍️ Author: Hiren Dave 📖 ℹ️ INFO: In the emerging field of private inference, each millisecond of latency directly translates into higher monetary cost, reduced user experience, and in some regulatory regimes, a breach of...... 🕒 Published: 2025-11-20 📂 Category: AI/ML 🔗 Read more: https://lnkd.in/duYktqZ8 🚀✨ #privateinference #latencyoptimizatio #coordinatedescent
-
AIxpanse
57 followers
Our inference latency jumped from 180ms to 2.1 seconds after we moved from 8 attention heads to 16 in our summarization model. Turns out we'd been misunderstanding how multi-head attention actually allocates compute, and it cost us three weeks of debugging. What was visible Prometheus showed P99 latency spiking hard. Users were complaining about timeouts on document summarization. We could see GPU utilization was fine at 60%, memory wasn't the bottleneck, but something was choking the forward pass in our vLLM deployment. 🐛 The investigation First thought was batch size. We tuned it down from 32 to 16, then 8. No improvement. Then we suspected the tokenizer was slow, added caching, profiled with py-spy. Still nothing. Spent two days convinced it was a vLLM version issue and rolled back. 🎯 What we missed Multi-head attention doesn't just split your dmodel across heads. Each head does a full QKV projection, attention computation, then concatenates. When we doubled heads from 8 to 16, we doubled the number of separate matrix multiplications happening sequentially on our A10 GPUs. The real kicker: our head dimension (dk) dropped from 64 to 32, which meant we lost the efficiency of larger GEMM operations. Small matrix ops are terrible for GPU throughput. How we resolved it Reverted to 8 heads but increased d_model from 512 to 768 to keep total parameters similar. Latency dropped back to 210ms. We also added explicit head dimension constraints in our model config validation so this can't happen silently again. 📚 What I do differently now I validate head dimensions during model init, not just total param count. Apparently reading the Vaswani paper once in 2017 doesn't mean you remember the compute implications in 2024. Config files will lie to you faster than any junior engineer ever could. #MultiHeadAttention #MLOps #TransformerModels #ProductionML
-
Alluxio
5K followers
GPU idle time is often a symptom of data architecture, not compute capacity. Coupang's ML platform faced a strict scheduling constraint: training jobs were physically tied to the specific clusters where the datasets resided. Migrating data to overflow GPU clusters during peak times required a manual, time-consuming data preparation step before jobs could even enter the queue. To solve this, they introduced a distributed cache to decouple compute and storage. ▪️ Decoupled Scheduling: Workloads can now run on any available GPU across multiple regions without requiring manual data migration first. ▪️ Continuous Pipelining: The manual data-copying step was eliminated. Jobs start immediately and pull data directly from the central data lake as needed. ▪️ Improved I/O: Utilizing local NVMe for the caching layer delivered ~40% better I/O performance compared to their standard cloud parallel file system. Read more: https://lnkd.in/gr26EDvs #MLOps #AIInfrastructure #GPU #DistributedSystems #PlatformEngineering
2
-
NimbleEdge
3K followers
We are very excited to publish our research on Contextual Sparsity in the PyTorch community. Explore our in-depth white paper https://lnkd.in/gcFN4abk for exploring what’s state of the art in Sparsity and how NimbleEdge is pushing the boundaries of on-device LLM inference optimization.
2
-
PyTorch
329K followers
Announcing torchcomms – a new PyTorch API for distributed programming designed to support scalability, fault tolerance, heterogeneity, and extensibility. Torchcomms makes it easier for developers to leverage innovations in existing and new collective communications backends like NCCL, RCCL, and NCCLX. 🖇️ Read the full blog from Team torchcomms at Meta here: https://lnkd.in/gdwbnMSV 👥 Contributors include: 🍃 Tristan Rice, Subodh Iyengar, Junjie W., Feng Tian, James Hongyi Zeng, Sudharssun Subramanian, Pavan Balaji, Qiye Tan, Rodrigo De Castro, Saif Hasan, Min Si, Yifan Mao, Dingming Wu, Zhaoyang Han, Blake Matheny, Art Zhu, Denis Boyda, Regina Ren, Jingyi Yang, Bingzhe Liu, Shuqiang Zhang, Mingran Yang, Cen Zhao, Adi Gangidi, Ashmitha Jeevaraj Shetty, Bruce Wu, Ching-Hsiang Chu, Yulun Wang, Srinivas Vaidyanathan, Chris Gottbrath, Davide Italiano, Shashi Gandham, Omar Baldonado #Torchcomms #PyTorch #API #OpenSourceAI #PyTorchCon
169
3 Comments -
Bryan Kian Hsiang Low
National University of… • 4K followers
Given a single model, how do we improve an #LLM’s reasoning performance with limited resources 💻 and inference time ⌛️? Can a smaller 1.5B model outperform a 7B model without incurring long inference time from sequential queries? In the work of Wenyang Hu, Gregory Lau, See-Kiong Ng, Bryan Kian Hsiang Low et al., we introduce the framework called Dipper to create #LLMs ensembles from an optimized set of diverse reasoning prompts to improve performance. Dipper runs queries in parallel with a prompt optimization method inspired by Determinantal Point Processes (DPP), making it super fast ⏩️ and effective. Furthermore, Dipper can work with LLM APIs without model access 📦! With Dipper, we demonstrated how a small ensemble of just three 1.5B models can outperform a 7B model on a range of math and non-math reasoning tasks, while taking almost the same inference time and just < 3x compute for a normal query thanks to accelerated batch inference methods 😱 ! Find out more at #EMNLP2025 (Poster 4492) at Hall C, Session 11, Nov 6 at 16:30! Paper: https://lnkd.in/gsDzJxPb
55
2 Comments -
Bunty Shah
MSCI Inc. • 4K followers
[AI Paper] We are finally treating Memory as an "Action Space," not just a "Storage Bucket." 🧠 As AI Architects, we usually treat Long-Term Memory (LTM) and Short-Term Memory (STM) as separate engineering problems. We use Vector DBs for LTM and Context Window sliding for STM. The agent has no real control over what it remembers; it's just passive retrieval. A new paper from Alibaba Group & Wuhan University, "Agentic Memory (AgeMem)," proposes a unified framework that changes this paradigm. The Architectural Shift: Memory as a Tool: Instead of a hidden background process, memory operations (Store, Retrieve, Update, Forget) are exposed as Tool Calls in the agent's action space. The model decides when to write to its own disk. Unified Policy: They don't use a separate "Memory Controller" model. They train the main agent (via a 3-stage RL process) to manage its own context, balancing LTM and STM dynamically. Progressive RL: The training recipe is key. It starts with "Memory-Free" warmups, moves to "Passive Memory" (learning to read), and ends with "Active Memory" (learning to write/manage). The Result: It outperforms standard RAG and heuristically managed memory systems (like MemGPT) on long-horizon tasks (HotpotQA) because the agent learns what is worth remembering. If you are building autonomous agents that need to run for days or weeks, "Passive Memory" isn't enough. You need "Active Memory Management." 👇 Link to the paper in the comments. #AIArchitecture #AgenticAI #MemorySystems #ReinforcementLearning #LLM #Alibaba #Research #AgeMem
18
1 Comment
Explore collaborative articles
We’re unlocking community knowledge in a new way. Experts add insights directly into each article, started with the help of AI.
Explore More