home/categories/containers/bagelhole-devops-security-agent-skills-infrastructure-local-ai-llm-inference-scaling-skill-md
containersdevops
llm-inference-scaling
Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. Handle traffic spikes, implement queue-based scaling, and optimize cost with spot instances for AI workloads.
maintainer
BagelHole
اپ ڈیٹ ہوا 3/2/2026
اسٹارز
18
فورکس
2
quick start
Installation and usage
Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. Handle traffic spikes, implement queue-based scaling, and optimize cost with spot instances for AI workloads.
انسٹالیشن
$ install --globalskills.sh
استعمال
انسٹال کرنے کے بعد، آپ یہ اسکل ٹرمینل میں درج ذیل کمانڈ چلا کر استعمال کر سکتے ہیں:
skills use llm-inference-scaling