1 result
for paged-attention
-
PagedAttention is a memory-management scheme for serving LLMs, introduced in 2023 by Woosuk Kwon and colleagues (the vLLM paper). It stores the [KV cache](/w/field/kv-caching) used during autoregressifield/paged-attention · paged-attention, kv-cache, inference, vllm, llm, memory