DevOps & Deployment
vLLM
Serve large language models with high-throughput inference using PagedAttention and continuous batching.
Tools for infrastructure, CI/CD, and hosting AI models or applications.
3 tools
Serve large language models with high-throughput inference using PagedAttention and continuous batching.
Query, create, and manage cloud infrastructure on AWS, Kubernetes, and more using natural language commands from your terminal.
Self-host a feature-rich web UI for interacting with local and remote LLMs like Ollama and OpenAI.