本课程由Vimal Daga主讲,深入讲解如何在Red Hat OpenShift AI平台上使用vLLM和KServe构建私有云大模型部署方案,涵盖Operator安装、资源调配、模型量化及故障处理,助你掌握企业级LLM服务实战技能。

原始标题:LLM on OpenShift AI: Deployment Masterclass

LLM on OpenShift AI: Deployment Masterclass

本课程由 Vimal Daga 主讲,聚焦 Red Hat OpenShift AI 平台,实战演示如何结合 vLLM 与 KServe 构建私有云企业级大语言模型部署方案。课程深入讲解了 Operator 安装、硬件资源调配、模型量化以及解决 OOM 等实战故障,助力实现高效的模型管理与应用对接。如需深入了解,建议在该平台上探索相关的开源文档与实战案例。

Published 8/2026
Created by Vimal Daga
MP4 | Video: h264, 1920×1080 | Audio: AAC, 44.1 KHz, 2 Ch
Level: All Levels | Genre: eLearning | Language: English | Duration: 8 Lectures ( 2h 49m ) | Size: 2.1 GB

Hands-On LLM Serving with OpenShift AI, vLLM, KServe & APIs

What you’ll learn
⚡ Understand the fundamentals of Large Language Models (LLMs), LLM inference, model runtimes, and private/sovereign AI architectures.
⚡ Understand how OpenShift AI is used to build, manage, and operate AI and LLM workloads on an OpenShift cluster.
⚡ Set up OpenShift AI infrastructure, install the OpenShift AI Operator, and configure a Data Science Cluster with components such as KServe, Workbench, Dashboard
⚡ Select appropriate LLM models based on model parameters, use cases, hardware requirements, and resource availability.
⚡ Understand LLM quantization and different model precision concepts and how they affect model size, memory consumption, performance, and deployment requirements.
⚡ Deploy and serve LLM models on OpenShift AI using technologies such as vLLM and KServe.
⚡ Configure CPU, memory requests, memory limits, hardware profiles, replicas, services, and OpenShift Routes for LLM workloads.
⚡ Troubleshoot common LLM deployment problems, including pending pods, insufficient cluster resources, CrashLoopBackOff, and out-of-memory errors using OpenShift
⚡ Test and interact with deployed LLM models using the OpenShift AI Playground and inference APIs

Requirements
❗ Familiarity with using a web browser and terminal/command-line interface.

Description
Generative AI is rapidly changing how modern applications are built, and Large Language Models (LLMs) are becoming an important part of enterprise AI infrastructure. This course is designed to give you apractical, hands-on understanding of deploying and serving LLMs using Red Hat OpenShift AI.

The course starts with the fundamentals ofLarge Language Models, LLM inference, model runtimes, model selection, and private or sovereign AI architectures. You will understand what happens behind the scenes when a user sends a prompt to an AI model and how an end-to-end LLM application can be built using your own infrastructure.

You will then move into theOpenShift AI platform, where you will learn how to prepare the infrastructure, install the OpenShift AI Operator, configure a Data Science Cluster, and understand important AI platform components such as KServe, Workbench, Dashboard, Model Registry, and AI pipelines.

The course also coversLLM model selection and optimization, including model parameters, memory requirements, model precision, and quantization. You will understand why different model formats require different infrastructure resources and how quantized models can help reduce memory requirements.

The core of the course is hands-onLLM deployment using OpenShift AI, vLLM, and KServe. You will configure hardware profiles, CPU and memory requests, resource limits, replicas, deployments, services, and OpenShift Routes while deploying an LLM for inference.

You will also learn how to troubleshoot real deployment problems, includingpending pods, insufficient CPU or memory, CrashLoopBackOff, and out-of-memory errors. The course demonstrates how to use OpenShift events, pod status, and container logs to identify and resolve issues.

Finally, you will work withLLM inference APIs and the OpenShift AI Playground, and learn how a deployed model can be integrated into a Python application usingGradio to create a ChatGPT-style web interface.

By the end of this course, you will have practical knowledge of the complete LLM deployment workflow—fromOpenShift AI infrastructure and model selection to model serving, troubleshooting, inference APIs, and application integration.

This course is ideal for DevOps Engineers, Cloud Engineers, Kubernetes and OpenShift professionals, MLOps Engineers, AI/ML Engineers, Platform Engineers, developers, and technology professionals who want to build practical skills inenterprise LLM deployment and AI infrastructure.

Who this course is for
⭐ DevOps and Cloud professionals looking to move into AI/ML workloads.
⭐ Kubernetes and OpenShift professionals interested in Generative AI and LLM infrastructure.
⭐ MLOps and AIOps professionals who want to understand AI model deployment and serving.
⭐ AI/ML engineers who want to learn the operational side of deploying LLMs.
⭐ Platform engineers and system administrators responsible for building AI platforms.
⭐ Developers who want to connect applications to privately hosted LLM inference APIs.
⭐ Students and technology enthusiasts who want practical exposure to enterprise LLM deployment.
⭐ Professionals interested in private, self-hosted, or sovereign AI infrastructure.

隐藏内容

此处内容需要权限查看

  • 普通3金币
  • 会员免费
  • 永久会员免费推荐
会员免费查看

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注