---
title: LLM Recommendations
description: Selecting the right language model depends on accuracy, speed, and cost. This article outlines key recommendations for evaluating and choosing LLMs for Ozgar.ai.
---

[Skip to content](https://help.ozgar.ai/en/ozgar/llm-recommendations#main-content)

English

Show submenu for translations

[Get support](https://help.ozgar.ai/en/ozgar/kb-tickets/new?hsLang=en)

Open main navigation

Close main navigation

- English
  
  Show submenu for translations
- [Get support](https://help.ozgar.ai/en/ozgar/kb-tickets/new)

 How can we help you?

- There are no suggestions because the search field is empty.

1. [Ozgar Documentation](https://help.ozgar.ai/en/ozgar?hsLang=en)
2. [Setup & Infrastructure](https://help.ozgar.ai/en/ozgar/setup-infrastructure?hsLang=en)
3. [LLMs](https://help.ozgar.ai/en/ozgar/setup-infrastructure?hsLang=en#llms)

July 15, 2026

# LLM Recommendations

## Practical guidance and recommendations for choosing, and deploying large language models.

### Quick Start - Azure OpenAI

To get started quickly, [create deployments](https://help.ozgar.ai/en/ozgar/llm-setup-guide?hsLang=en) for the following models in Azure OpenAI.

- gpt-5-nano
- gpt-5.4-mini
- gpt-5.4
- text-embedding-3-large

Configure each deployment

- [with the highest available rate limits](https://help.ozgar.ai/en/ozgar/llm-setup-guide?hsLang=en),
- [apply a permissive content filtering policy](https://help.ozgar.ai/en/ozgar/llm-content-filtering?hsLang=en),
- [and choose the correct region for data processing](https://help.ozgar.ai/en/ozgar/data-residency-azure-openai?hsLang=en).

### LLMs in Ozgar.ai

Ozgar AI allows you to configure different AI models depending on the task being executed. This helps balance quality, performance, throughput, and cost across the platform.

Model Settings are organized into four areas:

| **Area** | **Purpose** | **Model slots** |
| --- | --- | --- |
| Knowledge | Document processing, retrieval, ranking, and structured knowledge extraction | Standard, Advanced, Nano |
| Assistant | Chat, agent workflows, source code analysis, graph analysis, and code execution support | Assistant Model, Source Code Agent Model, Graph Agent Model, Data Filter Model |
| Documentation | Documentation and tutorial generation | Documentation Model |
| Embeddings | Vector encoding for semantic search and retrieval | Embedding Model |

Each model slot has its own **Provider** and **Model** selector. Providers and models can be configured independently per role. This means a customer can use one provider for all slots, or mix providers across individual model roles when the deployment supports those providers.

In most environments, the recommended provider is **Azure OpenAI.** Find a list of the recommended OpenAI models for each section below.

Note that to use OpenAI models directly through the OpenAI API, your account must be on [API Usage Tier 4 or higher](https://developers.openai.com/api/docs/guides/rate-limits#usage-tiers). Ozgar.ai may not function as intended on API usage tiers below Tier 4.

### 1. Knowledge Models

Knowledge Models are used during knowledge extraction and processing, including:

- Source file analysis
- Entity extraction
- Relationship detection
- Semantic enrichment
- Codebase and system structure analysis
- Structured extraction from large volumes of content
- Retrieval and ranking support

Knowledge Models typically process large volumes of data, so cost efficiency, stability, predictable output quality, and throughput are important.

This category accounts for the highest token consumption in Ozgar.ai.

#### Knowledge model slots

| **Slot** | **Purpose** |
| --- | --- |
| Standard | Standard knowledge extraction and processing |
| Advanced | More advanced reasoning, code analysis, and complex knowledge extraction |
| Nano | Lightweight, high-volume, structured extraction |

#### Recommendations

| **Slot** | **Best Quality** | **Best Value** |
| --- | --- | --- |
| Standard | gpt-5-mini | gpt-4o-mini |
| Advanced | gpt-5.4 | gpt-4.1 |
| Nano | gpt-5.4-nano | gpt-4.1-nano |

### 2. Assistant Models Recommendations

Assistant Models are used in the Ozgar AI Chat Assistant, including:

- Interactive Q&A
- Code explanations
- Reasoning over the knowledge base
- Workflow and impact analysis
- User-facing technical assistance
- Source code analysis
- Graph and relationship analysis
- Assistant data filtering
- Agent and code execution support

Assistant Models require strong reasoning, conversational quality, and the ability to interpret retrieved knowledge accurately. Since users directly interact with assistant outputs, these models should prioritize answer quality and reliability.

#### Assistant model slots

| **Slot** | **Purpose** |
| --- | --- |
| Assistant Model | Main chat assistant model |
| Source Code Agent Model | Supports source-code-focused assistant analysis |
| Graph Agent Model | Supports graph, relationship, and dependency analysis |
| Data Filter Model | Filters and prepares retrieved data for assistant responses |

#### Recommendations

| **Slot** | **Best Quality** | **Best Value** |
| --- | --- | --- |
| Assistant Model | gpt-5.4 | gpt-4.1 |
| Source Code Agent Model | gpt-5.4 | gpt-4.1 |
| Graph Agent Model | gpt-5.4-mini | gpt-4.1 |
| Data Filter Model | gpt-5.4-mini | gpt-4.1-mini |

### 3. Documentation Models Recommendations

#### Documentation model slots

Documentation Models are used for generating:

- Documentation pages
- Functional descriptions
- Technical explanations
- Architecture summaries
- Visualizations and narrative summaries
- Documentation tutorials

Documentation Models directly impact the quality, readability, and consistency of generated documentation. These models should prioritize clarity, writing quality, reasoning, and reliability.

#### Documentation model slots

| **Slot** | **Purpose** |
| --- | --- |
| Documentation Model | Generates documentation pages, tutorials, explanations, and summaries |

#### Recommendations

| **Slot** | **Best Quality** | **Best Value** |
| --- | --- | --- |
| Documentation Model | gpt-5.4 | gpt-4.1 |

### 4. Embedding Models Recommendations

Embedding Models are used for:

- Semantic search
- Vector storage
- Similarity matching
- Retrieval augmentation, also known as RAG
- Finding relevant knowledge base content for assistant responses

Embedding Models directly impact search accuracy and retrieval quality.

#### Embedding model slots

| **Slot** | **Purpose** |
| --- | --- |
| Embedding Model | Generates vector embeddings for semantic search and retrieval |

#### Recommendations

| **Slot** | **Best Quality** | **Best Value** |
| --- | --- | --- |
| Embedding Model | text-embedding-3-large | text-embedding-3-small |

 

- [Setup & Infrastructure](https://help.ozgar.ai/en/ozgar/setup-infrastructure?hsLang=en#main-content)

    - [LLMs](https://help.ozgar.ai/en/ozgar/setup-infrastructure?hsLang=en#llms)

[![](https://help.ozgar.ai/hubfs/ozgar-logo-black.svg)](http://ozgar.ai)

Copyright © 2026, Ozgar Inc.