Collection of cookbooks for Podman Quadlets
You can not select more than 25 topics Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.
 
 
 
 

24 lines
1005 B

##
## vLLM environment file
##
## Injected into every vLLM container launched by llmsnap (via --env-file in
## config.yaml). Installed with 0600 root:root permissions as it holds a secret.
##
# Hugging Face token, used to download gated / private model weights.
HF_TOKEN=hf_replace_me
# Optional: speed up weight downloads (requires hf_transfer in the image).
#HF_HUB_ENABLE_HF_TRANSFER=1
# VLLM container image
#
# CPU inference image: the OFFICIAL vLLM CI build (multi-arch amd64/arm64), used
# by the shipped config.yaml / model Quadlets to validate the whole chain WITHOUT
# a GPU. Its ENTRYPOINT is `vllm serve`, so the model Quadlets only append flags.
#
# For a real GPU deployment, swap this for the CUDA image, whose tag is
# architecture-specific (`docker.io/vllm/vllm-openai:v0.24.0` on amd64,
# `:v0.24.0-aarch64` on arm64) — Podman cannot template it at generation time.
# See SPECS.md for the production (GPU) reference.
VLLM_IMAGE=public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.24.0