vLLM with quantization

AI
python -m vllm.entrypoints.openai.api_server --model <model> --quantization awq