Skip to content

fix(vllm): raise default VLLM_GPU_UTILIZATION 0.4 → 0.55

Jérôme Revillard requested to merge fix/vllm-gpu-util-0.55 into main

Higher GPU utilization for contextual retrieval's doc-level context (65536 max_model_len). 0.55 of 48GB = 26.4GB for vLLM. Updated compose, all.yml, README_GPU, ChoosingLLMs.

Merge request reports

Loading