llm inference · part two

What one conversation costs, and how many fit

Two, times the layers whose cache grows, times the key-value heads, times the head dimension, times the bytes. Read layer_types first: layers that only see a sliding window stop growing, and counting them doubles your answer.

per token
·
per conversation
·
conversations that fit
·