Lightweight LLM API gateway with OpenAI and Anthropic Messages support.
- OpenAI Compatible API -
/v1/chat/completions - Anthropic Messages API -
/v1/messages - Multi-Upstream Support - Round-robin load balancing
- Model Mapping - Map local model names to upstream models
- Rate Limiting - Per-upstream RPM limiting
- API Key Authentication
- Response Validation - Fixes tool_call format issues
# Build
go build -o llm-gateway .
# Run
./llm-gatewayEdit config.toml:
[server]
host = "0.0.0.0"
port = 3000
[[upstream]]
name = "nvidia"
base_url = "https://integrate.api.nvidia.com/v1"
api_key = "nv-xxx"
timeout = 120
[ratelimit]
enabled = true
requests_per_minute = 40
[auth]
enabled = true
api_keys = ["sk-test123"]| Endpoint | Description |
|---|---|
GET /health |
Health check |
POST /v1/chat/completions |
OpenAI Chat Completions |
POST /v1/responses |
OpenAI Responses API |
POST /v1/messages |
Anthropic Messages |
GET /v1/models |
List models |
Configure model mappings in config.toml:
[models]
auto_chat = "nvidia/nemotron-3-nano-30b-a3b"MIT