Registry / llm-agents / onnxruntime-genai

onnxruntime-genai

JSON →
library0.14.1pypypi✓ verified 84d ago

ONNX Runtime GenAI is a Python library that provides an easy, flexible, and performant way to run Generative AI models (Large Language Models and multi-modal models) on-device and in the cloud using ONNX Runtime. It encapsulates the complete generative AI loop, including pre- and post-processing, inference with ONNX Runtime, logits processing, search and sampling, and KV cache management. The library is actively developed, with version 0.13.1 released in April 2026, generally following a quarterly release cadence in line with the broader ONNX Runtime project.

pip install onnxruntime-genai
INSTALL
IMPORT
SIG · ONNXRUNTIME-GENAI
O
onnxruntime-genai
llm-agentspythonv0.14.1
Install
Import
Disk
Pass rate
0/ 10
Env Coverage0 / 10
glibc
3.93.13
musl
3.93.13
Install & Compatibility
Where this runs
tested against v0.11.4 · pip install
no network on importno background threads
Install × environment matrix
Each cell = how many times install + import succeeded across repeated harness runs. Partial = flaky.
glibc = Debian/Ubuntu slim · musl = Alpine Linux
musl
glibc
py 3.10
✕ build_error
8/12 runs
py 3.11
✕ build_error
8/12 runs
py 3.12
✕ build_error
8/12 runs
py 3.13
✕ build_error
8/12 runs
py 3.9
✕ build_error
8/12 runs
Code
Verified usage

Verified import paths — ran on the pinned version, not inferred.

onnxruntime_genai
import onnxruntime_genai as og

This quickstart demonstrates how to load a pre-optimized ONNX model (like Phi-3 Mini), tokenize an input prompt, and generate text using the `onnxruntime-genai` library. Before running the Python code, you must download an ONNX model, typically using `huggingface-cli` into a local directory. The example uses environment variables for the model path for flexibility.

import os import onnxruntime_genai as og # --- Prerequisite: Download a model --- # The following shell command downloads the Phi-3 Mini 4K Instruct ONNX model (CPU-INT4 quantized). # You will need to install huggingface_hub: pip install huggingface_hub # Run this command in your terminal before executing the Python code: # huggingface-cli download microsoft/Phi-3-mini-4k-instruct-onnx \ # --include cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4/* \ # --local-dir ./phi-3-mini-onnx model_path = os.environ.get('ONNX_MODEL_PATH', './phi-3-mini-onnx') try: # 1. Load the model model = og.Model(model_path) print(f"Loaded {model.type} on {model.device_type}") # 2. Create a tokenizer tokenizer = og.Tokenizer(model) # 3. Create generator parameters params = og.GeneratorParams(model) params.set_search_options(max_length=200, top_p=0.9, temperature=0.7) # 4. Encode initial prompt and append to generator prompt = "The capital of France is" input_tokens = tokenizer.encode(prompt) # 5. Create a generator instance generator = og.Generator(model, params) generator.append_tokens(input_tokens) print(f"Prompt: {prompt}") print("Generated text:", end="") # 6. Generate tokens one by one and decode for streaming output while not generator.is_done(): generator.generate_next_token() last_token = generator.get_sequence(0)[-1] print(tokenizer.decode([last_token]), end="", flush=True) print() # Get the full decoded sequence (optional, for non-streaming output) # output = tokenizer.decode(generator.get_sequence(0)) # print(f"\nFull output: {output}") except Exception as e: print(f"An error occurred: {e}") print(f"Please ensure the model is downloaded to '{model_path}' and all dependencies are installed.")
Debug
Known issues
breakingAPI changes in version 0.6.0 for 'chat mode' (continuation/continuous decoding) introduced a breaking change. The `GeneratorParams.input_ids` attribute and `generator.compute_logits()` method were replaced or made redundant.
fix
Replace `params.input_ids = input_tokens` with `generator.append_tokens(input_tokens)` after the generator object is created. Remove calls to `generator.compute_logits()`. For multi-turn conversations, create a loop and call `generator.append_tokens(new_prompt_tokens)` for each turn.
affects: <=0.5.2
gotchaONNX Runtime GenAI versions 0.4.0 and earlier were incompatible with `transformers` library version 4.45.0 and later when using the Model Builder tool, leading to `RuntimeError: [json.exception.type_error.302]` if `tokenizer_config.json` contained an array for the `model_input_names` field.
fix
Upgrade `onnxruntime-genai` to 0.5.0 or later, or downgrade `transformers` to a version lower than 4.45.0.
affects: 0.1.0 - 0.4.0
gotchaPre-built wheels for `onnxruntime-genai` currently do not support Python 3.13. Attempting to install may result in `ERROR: No matching distribution found`.
fix
Use Python versions 3.10, 3.11, or 3.12 until official support for 3.13 is released.
affects: All versions
gotchaExamples in the `main` branch of the GitHub repository may not be compatible with the latest stable PyPI release binaries due to ongoing development.
fix
For stability, use examples from the corresponding version tag (e.g., `v0.13.1` branch) if using pre-built binaries, or build `onnxruntime-genai` from source if using examples directly from the `main` branch.
affects: All versions
breakingModels from earlier Ryzen AI releases are not compatible with Ryzen AI 1.7 (which uses OGA v0.11.2 from v0.9.2.2).
fix
If upgrading to Ryzen AI 1.7, download the updated, compatible models.
affects: Ryzen AI 1.6.1 and earlier models
Errors
Common errors & fixes
ModuleNotFoundError: No module named 'onnxruntime_genai'
The package `onnxruntime-genai` is not installed in the active Python environment or the environment is not correctly activated (e.g., in a Jupyter Notebook with the wrong kernel).
fix
Ensure `pip install onnxruntime-genai` was run successfully in the correct virtual environment, or switch to the appropriate Python kernel in your IDE/notebook.
ImportError: DLL load failed while importing onnxruntime_genai: A dynamic link library (DLL) initialization routine failed.
This usually occurs in a Conda environment on Windows due to an outdated C++ runtime for Visual Studio.
fix
In your Conda environment, run: `conda install conda-forge::vs2015_runtime`.
DLL load failed while importing onnxruntime_genai
On Windows with CUDA, this error often means the `CUDA_PATH` environment variable is not correctly set after CUDA Toolkit installation.
fix
Ensure the `CUDA_PATH` system environment variable is set to the installation directory of your CUDA Toolkit.
ERROR: No matching distribution found for onnxruntime-genai
The currently used Python version (e.g., Python 3.13) does not have pre-built wheels available for `onnxruntime-genai` on PyPI.
fix
Use Python 3.10, 3.11, or 3.12, which have supported pre-built distributions.
RuntimeError: [json.exception.type_error.302] type must be string, but is array.
Incompatibility between `onnxruntime-genai` (versions <= 0.4.0) and `HuggingFace transformers` (versions >= 4.45.0) when `tokenizer_config.json` uses an array for `model_input_names`.
fix
Upgrade `onnxruntime-genai` to version 0.5.0 or newer, or downgrade `transformers` to a version prior to 4.45.0.
Upgrade
Version history
0.14.1latest on PyPI · released Jun 2, 2026
Audit
Dependencies
onnxruntimerequiredCore runtime dependency; separated from onnxruntime-genai since version 0.4.0.
numpyrequiredRequired for array manipulation, especially for model inputs/outputs.
huggingface_huboptionalOften used for downloading ONNX models via `huggingface-cli` for local inference.
Agent activity
18 hits · last 30 days
node
14
Amazon
1
OpenAI (training)
1
Resources
onnxruntime-genai — pip install onnxruntime-genai · libregistry