# GPU CLI > Run code on cloud GPUs from your terminal. Prefix any command with `gpu` and it runs on a remote GPU. ## Docs - [AI Agent Workflows](https://gpu-cli.sh/docs/ai-agent-skill): Use GPU CLI with coding agents and repo-local automation workflows - [Commands Reference](https://gpu-cli.sh/docs/commands): Current reference for public GPU CLI commands - [Configuration](https://gpu-cli.sh/docs/configuration): Current gpu.jsonc reference for GPU CLI - [Hosted Proxy Quickstart](https://gpu-cli.sh/docs/hosted-proxy): Create a shared OpenAI-compatible endpoint for your organization - [GPU CLI Documentation](https://gpu-cli.sh/docs): Run code on cloud GPUs by prefixing any command with 'gpu run' - [LLM Inference](https://gpu-cli.sh/docs/llm): Run Ollama, vLLM, or llama.cpp on GPU pods with either local proxy routing or org-scoped hosted publish - [Organizations & Service Accounts](https://gpu-cli.sh/docs/organizations): Use organizations, sub-accounts, and service accounts for shared GPU CLI workflows - [Quickstart](https://gpu-cli.sh/docs/quickstart): Install GPU CLI, authenticate, and run your first pod or LLM workflow - [Serverless Endpoints](https://gpu-cli.sh/docs/serverless): Deploy and manage RunPod Serverless endpoints with GPU CLI - [Troubleshooting](https://gpu-cli.sh/docs/troubleshooting): Common issues and current limitations when using GPU CLI