This project would investigate key-value(KV)-cache offloading in large language model (LLM) inference engines, with particular emphasis on data placement and storage efficiency. First, the project would use an emulator such as FEMU to characterize workload access patterns, including hot-key distributions, write intensity, and write amplification factor (WAF). Then, we would evaluate the potential …
Supervisors:
Pınar Tözün, Jens Birk Andersen
Semester: Fall 2026
Tags: SSDs, LLMs, kv-cache, modern storage