[go: up one dir, main page]

DEV Community

#pytorch

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Your quantized model got worse, and nothing told you

Your quantized model got worse, and nothing told you

3
Comments 2
5 min read
Demystifying Shape Mismatches: How to Debug and Fix Tensor Dimensions in PyTorch

Demystifying Shape Mismatches: How to Debug and Fix Tensor Dimensions in PyTorch

Comments
4 min read
I Trained a 6.4M-Parameter Transformer From Scratch to Talk About Recipes

I Trained a 6.4M-Parameter Transformer From Scratch to Talk About Recipes

1
Comments
5 min read
RoPE: How 2D Rotations Solved Transformer Long-Context

RoPE: How 2D Rotations Solved Transformer Long-Context

1
Comments
4 min read
Running Qwen3 Through the ExecuTorch MLX Delegate: Up to 4.52x Faster on M1 Max

Running Qwen3 Through the ExecuTorch MLX Delegate: Up to 4.52x Faster on M1 Max

Comments
7 min read
Testing PyTorch 2.13 MPS FlexAttention on M1 Max: Up to 7.83x Faster for Sparse Attention

Testing PyTorch 2.13 MPS FlexAttention on M1 Max: Up to 7.83x Faster for Sparse Attention

Comments
7 min read
What Does `unsqueeze` Do in PyTorch? (And Why Your Model Keeps Asking For It)

What Does `unsqueeze` Do in PyTorch? (And Why Your Model Keeps Asking For It)

Comments
7 min read
PyTorch Broadcasting Explained: The 3 Rules (and the Silent Bug That Bites Everyone)

PyTorch Broadcasting Explained: The 3 Rules (and the Silent Bug That Bites Everyone)

1
Comments
4 min read
Debugging a Python "Memory Leak" That Was Actually a Measurement Bug (ru_maxrss vs VmRSS)

Debugging a Python "Memory Leak" That Was Actually a Measurement Bug (ru_maxrss vs VmRSS)

Comments
4 min read
Classifier-free guidance above 7.5 oversaturated our product renders

Classifier-free guidance above 7.5 oversaturated our product renders

1
Comments
4 min read
Using the channels-last memory format reduced the latency of our conversation backbone by 22%

Using the channels-last memory format reduced the latency of our conversation backbone by 22%

1
Comments
4 min read
The SDXL VAE overflow that decoded black images in fp16

The SDXL VAE overflow that decoded black images in fp16

1
Comments
4 min read
Data Science Workload: Giới hạn RAM trên Dell Pro Max 14 MC14250

Data Science Workload: Giới hạn RAM trên Dell Pro Max 14 MC14250

Comments
3 min read
Why My Medical AI Took 6.4 Seconds Per Scan and How I Got It to 3.1.

Why My Medical AI Took 6.4 Seconds Per Scan and How I Got It to 3.1.

1
Comments 2
6 min read
The seam our tiled upscaler left on every 4K product render

The seam our tiled upscaler left on every 4K product render

Comments
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.