[go: up one dir, main page]

DEV Community

#gpu

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Why Does Your Local Model Crash at 32k Tokens?

Why Does Your Local Model Crash at 32k Tokens?

Comments 1
12 min read
Same Source, Two Cost Curves: Local Preview vs GPU Final

Same Source, Two Cost Curves: Local Preview vs GPU Final

Comments
5 min read
Let your AI agent rent a GPU: llms.txt, --json and --budget

Let your AI agent rent a GPU: llms.txt, --json and --budget

1
Comments 1
4 min read
Evaluating Multi-Node LLM Orchestrators: Exo, GPUStack, and LocalAI

Evaluating Multi-Node LLM Orchestrators: Exo, GPUStack, and LocalAI

Comments
1 min read
How Much VRAM Do You Really Need to Run a 70B LLM?

How Much VRAM Do You Really Need to Run a 70B LLM?

Comments
7 min read
Same nvJPEG2000, different numbers: timer boundaries and frames in flight

Same nvJPEG2000, different numbers: timer boundaries and frames in flight

Comments
13 min read
Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture

Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture

1
Comments
10 min read
Why GPU Availability Is Still the Biggest Bottleneck in ML Infra

Why GPU Availability Is Still the Biggest Bottleneck in ML Infra

1
Comments 1
3 min read
DeepSpeed ZeRO vs PyTorch FSDP2, Measured: 1.8x on NVLink, Barely Finishing on PCIe (Plus How to Actually Get GPUs on GKE)

DeepSpeed ZeRO vs PyTorch FSDP2, Measured: 1.8x on NVLink, Barely Finishing on PCIe (Plus How to Actually Get GPUs on GKE)

Comments
10 min read
Half of my FP8 convolutions were silently running in f32 – a one-line XLA fix

Half of my FP8 convolutions were silently running in f32 – a one-line XLA fix

2
Comments 1
6 min read
What Happens When You Ask an LLM a Question

What Happens When You Ask an LLM a Question

Comments
8 min read
CUDA Cores vs Tensor Cores Explained

CUDA Cores vs Tensor Cores Explained

Comments 1
2 min read
Inside SFPU Overflow Bugs: How a 40-Year-Old Rounding Trick Breaks on Modern AI Accelerators

Inside SFPU Overflow Bugs: How a 40-Year-Old Rounding Trick Breaks on Modern AI Accelerators

Comments
3 min read
I Put a Paid AI Video Generator on My Own Gaming GPU — No Cloud Bill

I Put a Paid AI Video Generator on My Own Gaming GPU — No Cloud Bill

Comments 1
3 min read
Reverse-Engineering NVIDIA: Modifying a CUDA binary

Reverse-Engineering NVIDIA: Modifying a CUDA binary

Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.