<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>OME</title><link>https://ome-projects.github.io/ome/</link><description>Recent content on OME</description><generator>Hugo</generator><language>en</language><atom:link href="https://ome-projects.github.io/ome/index.xml" rel="self" type="application/rss+xml"/><item><title>OME API</title><link>https://ome-projects.github.io/ome/docs/reference/ome.v1beta1/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/reference/ome.v1beta1/</guid><description>&lt;p&gt;Package v1beta1 contains API Schema definitions for the serving v1beta1 API group&lt;/p&gt;
&lt;h2 id="resource-types"&gt;
Resource Types
&lt;a href="#resource-types" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://ome-projects.github.io/ome/docs/reference/ome.v1beta1/#ome-io-v1beta1-BaseModel"&gt;BaseModel&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ome-projects.github.io/ome/docs/reference/ome.v1beta1/#ome-io-v1beta1-BenchmarkJob"&gt;BenchmarkJob&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ome-projects.github.io/ome/docs/reference/ome.v1beta1/#ome-io-v1beta1-ClusterBaseModel"&gt;ClusterBaseModel&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ome-projects.github.io/ome/docs/reference/ome.v1beta1/#ome-io-v1beta1-ClusterServingRuntime"&gt;ClusterServingRuntime&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ome-projects.github.io/ome/docs/reference/ome.v1beta1/#ome-io-v1beta1-FineTunedWeight"&gt;FineTunedWeight&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ome-projects.github.io/ome/docs/reference/ome.v1beta1/#ome-io-v1beta1-InferenceService"&gt;InferenceService&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ome-projects.github.io/ome/docs/reference/ome.v1beta1/#ome-io-v1beta1-ServingRuntime"&gt;ServingRuntime&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="ome-io-v1beta1-BaseModel"&gt;
&lt;code&gt;BaseModel&lt;/code&gt;
&lt;a href="#ome-io-v1beta1-BaseModel" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Appears in:&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Labels and Annotations</title><link>https://ome-projects.github.io/ome/docs/reference/labels-and-annotations/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/reference/labels-and-annotations/</guid><description>&lt;p&gt;This document serves as a reference of the various labels and annotations used throughout OME.&lt;/p&gt;
&lt;h2 id="annotations"&gt;
Annotations
&lt;a href="#annotations" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;h3 id="inferenceservice-annotations"&gt;
InferenceService Annotations
&lt;a href="#inferenceservice-annotations" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;These annotations are used to configure InferenceService behavior:&lt;/p&gt;</description></item><item><title>Base Model</title><link>https://ome-projects.github.io/ome/docs/concepts/base_model/</link><pubDate>Tue, 14 Mar 2023 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/concepts/base_model/</guid><description>&lt;h2 id="what-is-a-base-model"&gt;
What is a Base Model?
&lt;a href="#what-is-a-base-model" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A Base Model in OME is a Kubernetes resource that represents a foundation AI model (like GPT, Llama, or Mistral) that you want to use for inference workloads. Think of it as a blueprint that tells OME where to find your model, how to download it, and where to store it on your cluster nodes.&lt;/p&gt;</description></item><item><title>Contributing to OME</title><link>https://ome-projects.github.io/ome/docs/developer-guide/contributing/</link><pubDate>Tue, 14 Mar 2023 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/developer-guide/contributing/</guid><description>&lt;p&gt;Thank you for your interest in contributing to OME! This repository is open to everyone and welcomes all kinds of contributions, no matter how small or large.&lt;/p&gt;
&lt;h2 id="ways-to-contribute"&gt;
Ways to Contribute
&lt;a href="#ways-to-contribute" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;There are several ways you can contribute to the project:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Bug Reports&lt;/strong&gt;: Identify and report issues or bugs&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Feature Requests&lt;/strong&gt;: Suggest new features or improvements&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Code Contributions&lt;/strong&gt;: Submit bug fixes or implement new features&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Documentation&lt;/strong&gt;: Improve documentation and guides&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Testing&lt;/strong&gt;: Help with testing and quality assurance&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="development-environment-setup"&gt;
Development Environment Setup
&lt;a href="#development-environment-setup" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;h3 id="prerequisites"&gt;
Prerequisites
&lt;a href="#prerequisites" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Before you begin, you&amp;rsquo;ll need to install these tools:&lt;/p&gt;</description></item><item><title>Deploy a Simple Inference Service</title><link>https://ome-projects.github.io/ome/docs/tasks/run-workloads/deploy-inference-service/</link><pubDate>Tue, 14 Mar 2023 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/tasks/run-workloads/deploy-inference-service/</guid><description>&lt;p&gt;This page shows you how to deploy a simple inference service using OME. You&amp;rsquo;ll learn how to create an InferenceService that serves a pre-trained model for real-time inference using SGLang and OpenAI-compatible APIs.&lt;/p&gt;
&lt;h2 id="before-you-begin"&gt;
Before you begin
&lt;a href="#before-you-begin" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;You need to have the following:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A Kubernetes cluster with OME installed&lt;/li&gt;
&lt;li&gt;&lt;code&gt;kubectl&lt;/code&gt; configured to communicate with your cluster&lt;/li&gt;
&lt;li&gt;GPU nodes available in your cluster (A100, H100, H200, or B4)&lt;/li&gt;
&lt;li&gt;Access to OME container registry (&lt;code&gt;ghcr.io/sgl-project/&lt;/code&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="step-1-verify-prerequisites"&gt;
Step 1: Verify prerequisites
&lt;a href="#step-1-verify-prerequisites" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Check that OME is installed and running:&lt;/p&gt;</description></item><item><title>Model Agent Administration</title><link>https://ome-projects.github.io/ome/docs/administration/model-agent/</link><pubDate>Tue, 14 Mar 2023 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/administration/model-agent/</guid><description>&lt;p&gt;The Model Agent is the core component responsible for downloading, managing, and distributing models across your OME cluster. This guide provides comprehensive information for cluster administrators who need to configure, monitor, and troubleshoot the Model Agent in production environments.&lt;/p&gt;
&lt;h2 id="architecture-overview"&gt;
Architecture Overview
&lt;a href="#architecture-overview" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;h3 id="daemonset-deployment"&gt;
DaemonSet Deployment
&lt;a href="#daemonset-deployment" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;The Model Agent is deployed as a Kubernetes DaemonSet, ensuring it runs on every node in your cluster. This distributed architecture provides several benefits:&lt;/p&gt;</description></item><item><title>Controller Configuration</title><link>https://ome-projects.github.io/ome/docs/administration/controller-configuration/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/administration/controller-configuration/</guid><description>&lt;p&gt;The OME controller manager (&lt;code&gt;ome-manager&lt;/code&gt;) is configured through command-line flags passed to its container. This page documents those flags and how to change them on a running cluster.&lt;/p&gt;
&lt;h2 id="setting-flags"&gt;
Setting flags
&lt;a href="#setting-flags" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The flags are parsed once at startup, so a change takes effect when the controller pod restarts. No image rebuild is required.&lt;/p&gt;</description></item><item><title>Runtime Revisions and Pinning</title><link>https://ome-projects.github.io/ome/docs/concepts/runtime-revision/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/concepts/runtime-revision/</guid><description>&lt;p&gt;By default an InferenceService always renders its pods from the &lt;strong&gt;live&lt;/strong&gt; ServingRuntime it references: whenever the runtime changes, the next reconcile picks up the new spec. Runtime &lt;strong&gt;pinning&lt;/strong&gt; lets you decouple an InferenceService from live runtime changes by pinning it to an immutable &lt;strong&gt;snapshot&lt;/strong&gt; of the runtime, so updates roll out only when you ask for them.&lt;/p&gt;
&lt;p&gt;Snapshots are stored as Kubernetes &lt;code&gt;ControllerRevision&lt;/code&gt; objects in the OME namespace, are content-addressed (deduplicated by hash), and are garbage-collected on a configurable schedule.&lt;/p&gt;</description></item><item><title>Run Performance Benchmarks</title><link>https://ome-projects.github.io/ome/docs/tasks/run-workloads/run-benchmarks/</link><pubDate>Tue, 14 Mar 2023 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/tasks/run-workloads/run-benchmarks/</guid><description>&lt;p&gt;This page shows you how to run performance benchmarks on your inference services using OME&amp;rsquo;s BenchmarkJob. You&amp;rsquo;ll learn how to test different traffic scenarios, measure performance metrics, and store results for analysis.&lt;/p&gt;
&lt;h2 id="before-you-begin"&gt;
Before you begin
&lt;a href="#before-you-begin" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;You need to have the following:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A Kubernetes cluster with OME installed&lt;/li&gt;
&lt;li&gt;&lt;code&gt;kubectl&lt;/code&gt; configured to communicate with your cluster&lt;/li&gt;
&lt;li&gt;An InferenceService deployed and ready&lt;/li&gt;
&lt;li&gt;Access to storage for benchmark results (OCI Object Storage or PVC)&lt;/li&gt;
&lt;li&gt;OME benchmark tool image available&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="step-1-verify-prerequisites"&gt;
Step 1: Verify prerequisites
&lt;a href="#step-1-verify-prerequisites" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Check that your inference service is running:&lt;/p&gt;</description></item><item><title>Fine Tuned Weight</title><link>https://ome-projects.github.io/ome/docs/concepts/fine_tuned_weight/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/concepts/fine_tuned_weight/</guid><description>&lt;h2 id="what-is-a-fine-tuned-weight"&gt;
What is a Fine Tuned Weight?
&lt;a href="#what-is-a-fine-tuned-weight" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A FineTunedWeight in OME is a Kubernetes resource that represents a set of weights that were fine-tuned from an existing &lt;a href="https://ome-projects.github.io/ome/docs/concepts/base_model"&gt;Base Model&lt;/a&gt;. Rather than describing a full, standalone model, it points back to a base model and carries only the artifacts and metadata specific to the fine-tuning - for example a LoRA adapter, an added adapter module, or a distilled variant.&lt;/p&gt;</description></item><item><title>Serving Runtime</title><link>https://ome-projects.github.io/ome/docs/concepts/serving_runtime/</link><pubDate>Tue, 14 Mar 2023 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/concepts/serving_runtime/</guid><description>&lt;p&gt;The only difference between the two is that one is namespace-scoped and the other is cluster-scoped.&lt;/p&gt;
&lt;p&gt;A &lt;em&gt;ClusterServingRuntime&lt;/em&gt; defines the templates for Pods that can serve one or more particular model.
Each ClusterServingRuntime defines key information such as the container image of the runtime and a list of the models that the runtime supports.
Other configuration settings for the runtime can be conveyed through environment variables in the container specification.&lt;/p&gt;</description></item><item><title>Inference Service</title><link>https://ome-projects.github.io/ome/docs/concepts/inference_service/</link><pubDate>Tue, 14 Mar 2023 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/concepts/inference_service/</guid><description>&lt;h2 id="what-is-an-inferenceservice"&gt;
What is an InferenceService?
&lt;a href="#what-is-an-inferenceservice" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;An InferenceService is the central Kubernetes resource in OME that orchestrates the complete lifecycle of model serving. It acts as a declarative specification that describes how you want your AI models deployed, scaled, and served across your cluster.&lt;/p&gt;
&lt;p&gt;Think of InferenceService as the &amp;ldquo;deployment blueprint&amp;rdquo; for your AI workloads. It brings together models (defined by BaseModel/ClusterBaseModel), runtimes (defined by ServingRuntime/ClusterServingRuntime), and infrastructure configuration to create a complete serving solution.&lt;/p&gt;</description></item><item><title>Benchmark</title><link>https://ome-projects.github.io/ome/docs/concepts/benchmark/</link><pubDate>Tue, 14 Mar 2023 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/concepts/benchmark/</guid><description>&lt;p&gt;A &lt;em&gt;BenchmarkJob&lt;/em&gt; is a resource in OME that automates the performance benchmarking of inference service or OCI Generative AI Service endpoints. It allows you to evaluate model serving performance under various traffic patterns and load conditions.&lt;/p&gt;
&lt;p&gt;BenchmarkJob uses &lt;a href="https://docs.sglang.ai/genai-bench/"&gt;genai-bench&lt;/a&gt;, a comprehensive benchmarking tool for evaluating generative AI model serving systems. For detailed information about genai-bench features and capabilities, refer to the &lt;a href="https://docs.sglang.ai/genai-bench/"&gt;official genai-bench documentation&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="core-components"&gt;
Core Components
&lt;a href="#core-components" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A BenchmarkJob consists of several key components:&lt;/p&gt;</description></item><item><title>Ingress and External Access</title><link>https://ome-projects.github.io/ome/docs/concepts/ingress/</link><pubDate>Fri, 21 Jun 2024 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/concepts/ingress/</guid><description>&lt;h2 id="what-is-ingress-in-ome"&gt;
What is Ingress in OME?
&lt;a href="#what-is-ingress-in-ome" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Ingress in OME provides external access to your AI inference services running inside the Kubernetes cluster. When you deploy an InferenceService, OME automatically creates the appropriate ingress resources based on your deployment mode and cluster configuration, allowing external clients to make API calls to your models.&lt;/p&gt;</description></item><item><title>Ingress Administration</title><link>https://ome-projects.github.io/ome/docs/administration/ingress/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/docs/administration/ingress/</guid><description>&lt;p&gt;OME (Open Model Engine) provides flexible ingress configuration to support various Kubernetes ingress controllers and deployment scenarios. This guide covers the configuration options available to cluster administrators.&lt;/p&gt;
&lt;h2 id="supported-ingress-controllers"&gt;
Supported Ingress Controllers
&lt;a href="#supported-ingress-controllers" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;OME supports the following ingress solutions:&lt;/p&gt;
&lt;h3 id="standard-kubernetes-ingress-controllers"&gt;
Standard Kubernetes Ingress Controllers
&lt;a href="#standard-kubernetes-ingress-controllers" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;NGINX Ingress Controller&lt;/strong&gt; - &lt;a href="https://kubernetes.github.io/ingress-nginx/deploy/"&gt;Installation Guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Traefik&lt;/strong&gt; - &lt;a href="https://doc.traefik.io/traefik/getting-started/install-traefik/"&gt;Installation Guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Kong Ingress Controller&lt;/strong&gt; - &lt;a href="https://developer.konghq.com/kubernetes-ingress-controller/install/"&gt;Installation Guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;HAProxy Ingress&lt;/strong&gt; - &lt;a href="https://haproxy-ingress.github.io/docs/getting-started/"&gt;Installation Guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Contour&lt;/strong&gt; - &lt;a href="https://projectcontour.io/getting-started/"&gt;Installation Guide&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="service-mesh-ingress"&gt;
Service Mesh Ingress
&lt;a href="#service-mesh-ingress" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Istio&lt;/strong&gt; (VirtualService + Gateway) - &lt;a href="https://istio.io/latest/docs/setup/install/"&gt;Installation Guide&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="gateway-api"&gt;
Gateway API
&lt;a href="#gateway-api" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Gateway API Controllers&lt;/strong&gt; - &lt;a href="https://gateway-api.sigs.k8s.io/guides/"&gt;Installation Guide&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="ome-ingress-configuration"&gt;
OME Ingress Configuration
&lt;a href="#ome-ingress-configuration" class="anchor-link"&gt;
 &lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;

&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;OME ingress behavior is configured through the &lt;code&gt;inferenceservice-config&lt;/code&gt; ConfigMap in the OME controller namespace.&lt;/p&gt;</description></item><item><title>Search Results</title><link>https://ome-projects.github.io/ome/search/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://ome-projects.github.io/ome/search/</guid><description/></item></channel></rss>