work Efficient LLM Serving for RAG (Cache-Craft) Chunk-cache decomposition and reuse for faster prefill. Efficient Video Generation through Patch-Level Caching Reusing intermediate states across similar patches via rectified flow interpolation. HeTLM — Heterogeneity-Aware User Behavior Prediction Clusterwise training showing small LMs can better model subjective browsing behaviors. fun