🧠Research Project
KV Cache Steering for Controlling Frozen LLMs
Controlling frozen LLM behavior at inference time by steering the key-value cache with no fine-tuning required.
Details
Project snapshot
KV Cache Steering manipulates a small language model’s key-value cache at inference time to induce step-by-step reasoning typically seen only in much larger models without any fine-tuning. Conducted with the Fundamental AI Lab / VIS Lab, the work is now on arXiv and under review.