Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions docs/ocr.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,12 +5,14 @@
| Cost Tracking | ✅ |
| Logging | ✅ (Basic Logging not supported) |
| Load Balancing | ✅ |
| Supported Providers | `mistral`, `azure_ai`, `vertex_ai` |
| Supported Providers | `mistral`, `azure_ai`, `vertex_ai`, `cohere` |

:::tip

LiteLLM follows the [Mistral API request/response for the OCR API](https://docs.mistral.ai/capabilities/vision/#optical-character-recognition-ocr)

The Cohere Parse OCR integration supports image inputs only

:::

## **LiteLLM Python SDK Usage**
Expand Down Expand Up @@ -347,4 +349,4 @@ The response follows Mistral's OCR format with the following structure:
| Mistral AI | [Usage](#quick-start) |
| Azure AI | [Usage](../docs/providers/azure_ocr) |
| Vertex AI | [Usage](../docs/providers/vertex_ocr) |

| Cohere | [Usage](../docs/providers/cohere#parse-ocr) |
46 changes: 44 additions & 2 deletions docs/providers/azure_ocr.md
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
# Azure AI OCR (Mistral)
# Azure AI OCR

## Overview

Expand Down Expand Up @@ -146,9 +146,51 @@ response = await litellm.aocr(
Azure AI OCR endpoints don't have internet access. LiteLLM automatically converts public URLs to base64 data URIs before sending requests to Azure AI.
:::

## Cohere Parse

LiteLLM supports Cohere Parse v5 through Azure AI Foundry with the `azure_ai/Cohere-parse-v5` route. Set `AZURE_AI_API_KEY` and `AZURE_AI_API_BASE`, for example `https://<resource>.services.ai.azure.com/models`. LiteLLM rewrites this base URL to `/providers/cohere/v2/parse`

### LiteLLM SDK

```python showLineNumbers title="Cohere Parse SDK Usage"
import litellm
import os

os.environ["AZURE_AI_API_KEY"] = "your Azure AI key"
os.environ["AZURE_AI_API_BASE"] = "https://<resource>.services.ai.azure.com/models"

response = litellm.ocr(
model="azure_ai/Cohere-parse-v5",
document={
"type": "image_url",
"image_url": "https://example.com/image.png",
},
output_format="markdown",
)
```

### LiteLLM Proxy

```yaml showLineNumbers title="proxy_config.yaml"
model_list:
- model_name: cohere-parse-v5
litellm_params:
model: azure_ai/Cohere-parse-v5
api_key: "os.environ/AZURE_AI_API_KEY"
api_base: "os.environ/AZURE_AI_API_BASE"
model_info:
mode: ocr
```

The `azure_ai` provider fronts several OCR backends and picks one from the model name in `litellm_params.model`. Any name containing both `cohere` and `parse`, case-insensitively, routes to Cohere Parse, so a custom Foundry deployment name works as long as it contains both strings. A name missing either string falls through to the Mistral OCR route and fails against a Cohere endpoint, so keep `litellm_params.model` as `azure_ai/Cohere-parse-v5` if your deployment is named differently; the `model_name` alias can be anything

Cohere Parse accepts image inputs only. Use an `image_url` document with an image URL or a base64 image data URI. `document_url` and PDF inputs raise an error. Local image files passed with `{"type": "file", ...}` are converted to data URIs by LiteLLM

Foundry cannot fetch external image URLs, so LiteLLM downloads remote images and inlines them as base64 data URIs before sending the request. The response uses the standard OCR format with `pages[].index`, `pages[].markdown`, `pages[].images`, and `usage_info.pages_processed`

## Supported Models

- `mistral-document-ai-2505` - Latest Mistral OCR model on Azure AI
- `Cohere-parse-v5` - Cohere Parse v5 OCR model on Azure AI

Use the Azure AI provider prefix: `azure_ai/<model-name>`

68 changes: 68 additions & 0 deletions docs/providers/cohere.md
Original file line number Diff line number Diff line change
Expand Up @@ -207,6 +207,74 @@ print(response)
</TabItem>
</Tabs>

## Parse (OCR)

LiteLLM supports Cohere Parse v5 for image OCR through the direct Cohere route `cohere/parse-v5.0`. The direct route uses `COHERE_API_KEY`, sends requests to `https://api.cohere.com/v2/parse`, and is available through `litellm.ocr()`, `litellm.aocr()`, and the proxy OCR endpoint

<Tabs>
<TabItem value="sdk" label="LiteLLM SDK Usage">

```python
import litellm
import os

os.environ["COHERE_API_KEY"] = "cohere key"

response = litellm.ocr(
model="cohere/parse-v5.0",
document={
"type": "image_url",
"image_url": "https://example.com/image.png",
},
output_format="markdown",
)

for page in response.pages:
print(page.index)
print(page.markdown)
```

Set `output_format="blocks"` to include `pages[].blocks` in the standard `OCRResponse`. The default output format is `markdown`. Set `req_format="native"` to return the raw Cohere response

</TabItem>

<TabItem value="proxy" label="LiteLLM Proxy Usage">

Add this to your LiteLLM proxy config.yaml

```yaml
model_list:
- model_name: cohere-parse-v5
litellm_params:
model: cohere/parse-v5.0
api_key: os.environ/COHERE_API_KEY
model_info:
mode: ocr
```

Call the proxy at `/v1/ocr` or `/ocr`

```bash
curl http://0.0.0.0:4000/v1/ocr \
-H "Authorization: Bearer sk-1234" \
-H "Content-Type: application/json" \
-d '{
"model": "cohere-parse-v5",
"document": {
"type": "image_url",
"image_url": "https://example.com/image.png"
},
"output_format": "markdown"
}'
```

</TabItem>
</Tabs>

Cohere Parse accepts images only. Use an `image_url` document with an image URL or a base64 image data URI. `document_url` and PDF inputs raise an error. Local image files passed with `{"type": "file", ...}` are converted to data URIs by LiteLLM

The standard OCR response includes `pages[].index`, `pages[].markdown`, `pages[].images`, and `usage_info.pages_processed`. Cohere Parse costs $0.0015 per page


## Supported Models
| Model Name | Function Call |
Expand Down
Loading