Address
:
[go:
up one dir
,
main page
]
Include Form
Remove Scripts
Accept Cookies
Show Images
Show Referer
Rotate13
Base64
Strip Meta
Strip Title
Session Cookies
Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
cuda
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
How Fast Can a 421M-Parameter Decision Model Run? I Benchmarked Laya Across NVIDIA GPUs
Bhushan Kinge
Bhushan Kinge
Bhushan Kinge
Follow
Sep 24
How Fast Can a 421M-Parameter Decision Model Run? I Benchmarked Laya Across NVIDIA GPUs
#
machinelearning
#
cuda
#
performance
#
opensource
Comments
1
 comment
6 min read
Same nvJPEG2000, different numbers: timer boundaries and frames in flight
Fyodor Serzhenko
Fyodor Serzhenko
Fyodor Serzhenko
Follow
Sep 18
Same nvJPEG2000, different numbers: timer boundaries and frames in flight
#
cuda
#
gpu
#
performance
#
cpp
Comments
Add Comment
13 min read
Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture
ComradePenguin
ComradePenguin
ComradePenguin
Follow
Sep 18
Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture
#
cuda
#
nvidia
#
gpu
1
 reaction
Comments
Add Comment
10 min read
Nvidia Just Let Rust Into CUDA. Here's Why That's a Bigger Deal Than It Sounds
Ashraf
Ashraf
Ashraf
Follow
Sep 18
Nvidia Just Let Rust Into CUDA. Here's Why That's a Bigger Deal Than It Sounds
#
rust
#
ai
#
cuda
#
programming
1
 reaction
Comments
Add Comment
5 min read
A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 23
A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x
#
gemma
#
llamacpp
#
cuda
#
benchmarking
9
 reactions
Comments
2
 comments
11 min read
CUDA Rust: two native tracks, not a wrapper over C++
Juan Torchia
Juan Torchia
Juan Torchia
Follow
Sep 17
CUDA Rust: two native tracks, not a wrapper over C++
#
english
#
rust
#
cuda
#
gpuprogramming
1
 reaction
Comments
Add Comment
5 min read
Reverse-Engineering NVIDIA: Modifying a CUDA binary
Stjepan
Stjepan
Stjepan
Follow
Sep 6
Reverse-Engineering NVIDIA: Modifying a CUDA binary
#
nvidia
#
gpu
#
cuda
#
hex
Comments
Add Comment
4 min read
Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 22
Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It
#
gemma
#
vllm
#
gcp
#
cuda
13
 reactions
Comments
Add Comment
13 min read
Finding a Random Island with Geometry and CUDA
Muhammad Adil
Muhammad Adil
Muhammad Adil
Follow
Aug 19
Finding a Random Island with Geometry and CUDA
#
gpucomputing
#
geospatial
#
cuda
#
geometry
Comments
Add Comment
4 min read
Gemma 4 on a Tesla T4: QAT Weights Decode 1.79x Faster Than bf16
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 18
Gemma 4 on a Tesla T4: QAT Weights Decode 1.79x Faster Than bf16
#
gemma
#
vllm
#
cuda
#
machinelearning
8
 reactions
Comments
1
 comment
9 min read
Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU
Michał Piszczek
Michał Piszczek
Michał Piszczek
Follow
Aug 17
Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU
#
aiinfrastructure
#
llm
#
cuda
#
performance
1
 reaction
Comments
1
 comment
11 min read
Understanding GPU Memory: VRAM, Bandwidth, and Why Your Model Won't Fit
Aarush Karak
Aarush Karak
Aarush Karak
Follow
Aug 13
Understanding GPU Memory: VRAM, Bandwidth, and Why Your Model Won't Fit
#
gpu
#
cuda
#
vram
#
memory
1
 reaction
Comments
Add Comment
2 min read
Gemma 4 on an 2021 4 GB Laptop GPU: QAT Takes It From 9.5 GiB to 1.6
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 10
Gemma 4 on an 2021 4 GB Laptop GPU: QAT Takes It From 9.5 GiB to 1.6
#
gemma
#
llamacpp
#
mcp
#
cuda
11
 reactions
Comments
3
 comments
13 min read
The sm_120 shared-memory cliff: why FP8 KV cache crashes vLLM on workstation Blackwell
Jahn
Jahn
Jahn
Follow
Aug 25
The sm_120 shared-memory cliff: why FP8 KV cache crashes vLLM on workstation Blackwell
#
cuda
#
llm
#
gpu
#
performance
Comments
Add Comment
3 min read
g5g vs g6 for LLM Serving: the Same Code, and 3.7x the Throughput
xbill
xbill
xbill
Follow
for
AWS Community Builders
Aug 31
g5g vs g6 for LLM Serving: the Same Code, and 3.7x the Throughput
#
aws
#
jax
#
cuda
#
machinelearning
5
 reactions
Comments
2
 comments
6 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account