[go: up one dir, main page]

DEV Community

Mingxin Technology profile picture

Mingxin Technology

Storage acceleration for LLM inference. KV-cache tiering on NVMe-oF: +29-40% throughput (signed benchmarks, reproducible). mingxinstorage.xyz

Joined Joined on 
KV Cache Reuse in Multi-Turn Dialogue: A Deployment Case Study

KV Cache Reuse in Multi-Turn Dialogue: A Deployment Case Study

Comments
5 min read
How KV Cache Pooling and Sharing Improves Inference Resource Utilization

How KV Cache Pooling and Sharing Improves Inference Resource Utilization

Comments
5 min read
Distributed KV Cache Load Balancing: Strategies and Measured Trade-offs

Distributed KV Cache Load Balancing: Strategies and Measured Trade-offs

Comments
6 min read
Measured Evaluation of KV Cache Acceleration for Inference in Online Education Scenarios

Measured Evaluation of KV Cache Acceleration for Inference in Online Education Scenarios

Comments
4 min read
Deployment Path and Measured Performance of Domestic AI Inference Acceleration Cards in Real-Time Database Queries

Deployment Path and Measured Performance of Domestic AI Inference Acceleration Cards in Real-Time Database Queries

Comments
5 min read
Case Analysis of Huawei OceanStor UCM Inference Acceleration Solution

Case Analysis of Huawei OceanStor UCM Inference Acceleration Solution

Comments
5 min read
Why KV Cache Tiering Is the Key TCO Optimization Point in a Three-Tier Storage Architecture for Compute Centers

Why KV Cache Tiering Is the Key TCO Optimization Point in a Three-Tier Storage Architecture for Compute Centers

Comments
6 min read
The Hidden Cost of Model Switching: A Measured Path from 46.7% to 62.8% Effective Compute Utilization

The Hidden Cost of Model Switching: A Measured Path from 46.7% to 62.8% Effective Compute Utilization

Comments
5 min read
Cross-Platform Methodology Porting: Validating KV Cache Memory Efficiency from Muxi N260 to MI308X

Cross-Platform Methodology Porting: Validating KV Cache Memory Efficiency from Muxi N260 to MI308X

Comments
3 min read
Inference Deployment in Xinchuang Scenarios: Architecture Essentials for Data Not Leaving the Domain

Inference Deployment in Xinchuang Scenarios: Architecture Essentials for Data Not Leaving the Domain

Comments
5 min read
Clos Network Architecture: A Cost-Effectiveness and Selection Framework for Thousand-Card Inference Clusters

Clos Network Architecture: A Cost-Effectiveness and Selection Framework for Thousand-Card Inference Clusters

Comments
7 min read
How to Determine the Storage-to-Compute Ratio for Inference Clusters: Measured Basis for One Array Serving 8 Nodes

How to Determine the Storage-to-Compute Ratio for Inference Clusters: Measured Basis for One Array Serving 8 Nodes

Comments
6 min read
Engineering Practice Essentials of NVMe-oF + RoCEv2 in Inference Storage Scenarios: A Case Study with FX100 Measurements

Engineering Practice Essentials of NVMe-oF + RoCEv2 in Inference Storage Scenarios: A Case Study with FX100 Measurements

Comments
5 min read
What 90% Line-Rate Utilization on a Single 100GbE Port Means: Analyzing Network Bottlenecks in Inference Storage

What 90% Line-Rate Utilization on a Single 100GbE Port Means: Analyzing Network Bottlenecks in Inference Storage

Comments
5 min read
loading...