Scratch Pad

Thoughts on generative models, video quality, and the intersection of perception and computation.

Writing

6 posts

RL for Video: The State of the Space

A short field note on post-training for video generation in mid-2026: the pretraining recipe that converged, the physics and persistence failures scale is not fixing, the reward models and GRPO variants industry already runs, the three constraints that make video RL genuinely hard, and why it is still the right bet.

Read more →

CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion

Best-of-N search pays full price for every candidate it throws away. This post shows that the candidate rankings search depends on survive aggressive training-free caching, so you can explore every candidate cheaply and re-generate only the winner at full compute: 94.7% of best-of-8's gain at 63% of its cost, with any cache engine and any verifier.

Read more →

Fleetcraft: GPU Infrastructure Notes from Video Generation Projects

What three video-generation projects on TACC's Vista and Lonestar6 taught us about GPU infrastructure: memory hierarchy and OOM taxonomy, calibration-first batching, queue-driven multi-node fleets, asynchronous RL orchestration, telemetry that tells the truth, and the measured utilization evidence behind each lesson.

Read more →

Prepping for a Research Scientist, GenAI Position: A Pointer Notebook

A long revision notebook for Research Scientist / GenAI loops focused on image generation, perceptual quality, and video processing: diffusion and flow models, transformer internals (attention variants, RoPE, KV cache), the text-to-image design space, evaluation metrics, color and HDR, classical CV, RL alignment, and the coding tier, all in one place.

Read more →

Talks & Slides

4 decks

Press F or the button for fullscreen; use arrow keys to navigate.