8B on one H100
The comfortable case. 15 GiB of weights against a 72 GiB budget leaves 57 GiB of KV cache, which is 57 sequences at 8k or one at 467k.
- model
- llama-3.1-8b-instruct
- gpu-count
- 1
- gpu-vram
- 80
- max-model-len
- 8192
- max-num-seqs
- 256
- gpu-memory-utilization
- 0.9
- quant
- fp16
- kv-dtype
- auto
- prefix-caching
- yes
- host
- 127.0.0.1
- port
- 8000