Saved build
Saved AI stack
/b/b_build_005 · saved 2026-07-06
The catalog changed since this build was saved. The sticker shows current parts and prices; the saved estimated monthly was $1452.70/mo on the 24/7 basis.
Required parts
All required parts are selected for Cloud LLM self-hosted on a rented GPU.
Compatibility
?Compatibility needs review: 1 compatibility claim needs a manual check.
- Gemma 3 27B with vLLM has no current reviewed positive compatibility evidence in this catalog. Choose an exact model artifact and runtime Offering with reviewed support, or confirm their fit manually before deploying.
typical $1452.70–$1452.70
Technical evidence for selected parts
Host: Lambda A100 SXMEstimate · GPU rental · 24/7; the rental is the token costCatalog identity · Reviewed 69d ago · stale
Model: Gemma 3 27BEstimate · self-hosted — the compute cost carries the tokensCatalog identity · Reviewed 90d ago · stale
Model runner: vLLMCatalog identity · Reviewed 86d ago · aging
Interface: Command lineCatalog identity · Reviewed 95d ago · stale
Build estimate · saved 2026-07-06
Use this build
Print it, export its data, compare it, or make an editable copy. This saved link will not change.
Setup checklist
- Provision or prepare Host: Lambda A100 SXM [Catalog identity: Reviewed 69d ago · stale] [Estimate: GPU rental · 24/7; the rental is the token cost].
- Configure Model: Gemma 3 27B [Catalog identity: Reviewed 90d ago · stale] [Estimate: self-hosted — the compute cost carries the tokens].
- Configure Model runner: vLLM [Catalog identity: Reviewed 86d ago · aging].
- Connect Interface: Command line [Catalog identity: Reviewed 95d ago · stale].
- Recheck compatibility after changing any part.
- Review the build caveats before provisioning or purchasing anything.
- Run a small end-to-end test before moving the setup into regular use.
Things to know
- vLLM's command line launches and administers its inference server; it is not an interactive chat CLI. Reach a vLLM deployment through its API endpoint.