[go: up one dir, main page]

Hey, I'm Shubham.

I'm a machine learning engineer, about eight years in, mostly teaching machines to see and read: computer vision, multimodal models, and recommender systems, most recently at Meta on Instagram Ads ranking. Lately I'm most excited about multimodal LLMs, post-training, and evals.

The projects I love most fix a problem I actually ran into, like curbcheck, a model I trained to read SF parking signs after collecting two tickets in one week. When I'm not on something like that or artifold, I'm following Arsenal a little too closely, and splitting the rest between poker and a recent, slightly out-of-hand obsession with padel.

Currently going deep on multimodal post-training and evals. what I'm up to now →

Things I'm building

All projects
curbcheck screenshot

Can a small VLM tell you if you can park in San Francisco? A parking-sign benchmark plus QLoRA-tuned 3B and 7B students that read the pole; a deterministic resolver does the logic.

7B student reads real SF sign poles at 0.83 F1 (base: 0.22); verdicts right ~0.9 via the resolver

  • VLM
  • QLoRA
  • Qwen2.5-VL
  • evals
  • Modal
artifold screenshot

A local-first library for the stuff you make with AI. Index, search, preview, and share your work, then use your past output as the style guide for the next thing you build.

Published on PyPI, pip-installable

  • Python
  • CLI
  • embeddings
  • semantic search
rl-arcade screenshot

A gamified app that teaches reinforcement learning across six levels of PufferLib games, with interactive in-browser training, concept visualizations, and an XP system.

6 playable levels, train models live in the browser

  • Streamlit
  • RL
  • PufferLib

Writing

All posts

Say hi

Always happy to talk multimodal models, recsys, football, or whatever you're building.