Skip to content

Commit ceacd66

Browse files
authored
Merge pull request #14 from aprilgittens/main
Support Foundry OpenAI v1 and Responses API endpoints in live evaluations
2 parents 56d0243 + dc886e4 commit ceacd66

19 files changed

Lines changed: 375 additions & 151 deletions

‎.env.example‎

Lines changed: 6 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -1,21 +1,21 @@
11
# Copy this file to .env before running a live evaluation.
22
# .env is ignored by Git and must never be committed.
33

4-
# Azure Model Router endpoint
5-
AZURE_MODEL_ROUTER_ENDPOINT=https://your-resource.services.ai.azure.com/models
4+
# Foundry OpenAI v1 base URL (do not include /chat/completions)
5+
AZURE_MODEL_ROUTER_ENDPOINT=https://your-resource.services.ai.azure.com/openai/v1
66
AZURE_MODEL_ROUTER_KEY=your-model-router-api-key
77
AZURE_MODEL_ROUTER_DEPLOYMENT=model-router
88

9-
# Azure OpenAI endpoint (for baseline model)
10-
AZURE_OPENAI_ENDPOINT=https://your-resource.openai.azure.com
9+
# Foundry OpenAI v1 base URL for the baseline (do not include /responses)
10+
AZURE_OPENAI_ENDPOINT=https://your-resource.services.ai.azure.com/openai/v1
1111
AZURE_OPENAI_KEY=your-azure-openai-api-key
1212
# Must match a key in the pricing section of the config file you run (e.g. gpt-4o, gpt-5).
1313
# If you use a custom deployment name, add a matching entry under pricing: in that same config
1414
# using your exact deployment name. Otherwise baseline costs show as $0.00.
1515
AZURE_BASELINE_DEPLOYMENT=gpt-5
1616

17-
# Judge model (defaults to same endpoint as baseline; override to use a different model)
18-
AZURE_JUDGE_ENDPOINT=https://your-resource.openai.azure.com
17+
# Judge model (defaults to the same Foundry OpenAI v1 base URL as the baseline)
18+
AZURE_JUDGE_ENDPOINT=https://your-resource.services.ai.azure.com/openai/v1
1919
AZURE_JUDGE_KEY=your-azure-openai-api-key
2020
AZURE_JUDGE_DEPLOYMENT=gpt-5
2121

‎README.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -164,13 +164,13 @@ cp .env.example .env
164164

165165
| Variable | Description |
166166
|----------|-------------|
167-
| `AZURE_MODEL_ROUTER_ENDPOINT` | Model Router endpoint URL |
167+
| `AZURE_MODEL_ROUTER_ENDPOINT` | Foundry OpenAI v1 base URL ending in `/openai/v1` |
168168
| `AZURE_MODEL_ROUTER_KEY` | Model Router API key |
169169
| `AZURE_MODEL_ROUTER_DEPLOYMENT` | Model Router deployment name (e.g. `model-router`) |
170-
| `AZURE_OPENAI_ENDPOINT` | Azure OpenAI endpoint URL (baseline) |
170+
| `AZURE_OPENAI_ENDPOINT` | Baseline OpenAI v1 base URL ending in `/openai/v1` |
171171
| `AZURE_OPENAI_KEY` | Azure OpenAI API key (baseline) |
172172
| `AZURE_BASELINE_DEPLOYMENT` | Baseline model deployment name (e.g. `gpt-5`) |
173-
| `AZURE_JUDGE_ENDPOINT` | Judge model endpoint URL (can be same as baseline) |
173+
| `AZURE_JUDGE_ENDPOINT` | Judge OpenAI v1 base URL (can be same as baseline) |
174174
| `AZURE_JUDGE_KEY` | Judge model API key |
175175
| `AZURE_JUDGE_DEPLOYMENT` | Judge model deployment name (e.g. `gpt-5`) |
176176
| `AZURE_AI_PROJECT_ENDPOINT` | Microsoft Foundry project endpoint (optional, for cloud eval) |

‎configs/README.md‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -5,13 +5,18 @@
55
| [default.yaml](default.yaml) | All | Enabled | Standard evaluation |
66
| [quick_test.yaml](quick_test.yaml) | 5 | Disabled | Fast smoke test |
77
| [large_scale.yaml](large_scale.yaml) | All | Enabled | 1000+ prompts (higher concurrency, longer timeouts) |
8+
| [live_demo.yaml](live_demo.yaml) | All | Enabled | Foundry OpenAI v1 live demo |
89
| [foundry.yaml](foundry.yaml) | — | — | Foundry cloud eval settings (graders, thresholds) |
910

1011
Edit `default.yaml` to set your endpoint URLs, baseline model, and pricing.
1112
Environment variables (`${VAR}`) are resolved from `.env`. Supported Foundry
1213
model prices are refreshed from the Azure Retail Prices API and cached under
1314
`.cache/`; YAML prices are retained as offline and ambiguity fallbacks.
1415

16+
Foundry OpenAI v1 endpoints use `type: openai_compatible` and a base URL
17+
ending in `/openai/v1`. Set `api_mode` to `chat_completions` or `responses`;
18+
do not append the operation name to `endpoint_url`.
19+
1520
## Prompt Templates
1621

1722
- [`judge_prompts/`](judge_prompts/) — Local LLM-as-judge prompt templates (absolute and pairwise)

‎configs/default.yaml‎

Lines changed: 6 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -9,21 +9,21 @@ evaluation:
99

1010
endpoints:
1111
model_router:
12-
type: "azure_openai"
12+
type: "openai_compatible"
13+
api_mode: "chat_completions"
1314
endpoint_url: "${AZURE_MODEL_ROUTER_ENDPOINT}"
1415
api_key: "${AZURE_MODEL_ROUTER_KEY}"
1516
deployment_name: "${AZURE_MODEL_ROUTER_DEPLOYMENT}"
1617
parameters:
17-
temperature: 0.7
1818
max_tokens: 1024
1919

2020
baseline:
21-
type: "azure_openai"
21+
type: "openai_compatible"
22+
api_mode: "responses"
2223
endpoint_url: "${AZURE_OPENAI_ENDPOINT}"
2324
api_key: "${AZURE_OPENAI_KEY}"
2425
deployment_name: "${AZURE_BASELINE_DEPLOYMENT}"
2526
parameters:
26-
temperature: 0.7
2727
max_tokens: 1024
2828

2929
pricing: # USD per 1M tokens
@@ -172,7 +172,8 @@ output:
172172
judge:
173173
enabled: true
174174
endpoint:
175-
type: "azure_openai"
175+
type: "openai_compatible"
176+
api_mode: "responses"
176177
endpoint_url: "${AZURE_JUDGE_ENDPOINT}"
177178
api_key: "${AZURE_JUDGE_KEY}"
178179
deployment_name: "${AZURE_JUDGE_DEPLOYMENT}"

‎configs/large_scale.yaml‎

Lines changed: 6 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -10,15 +10,17 @@ evaluation:
1010

1111
endpoints:
1212
model_router:
13-
type: "azure_openai"
13+
type: "openai_compatible"
14+
api_mode: "chat_completions"
1415
endpoint_url: "${AZURE_MODEL_ROUTER_ENDPOINT}"
1516
api_key: "${AZURE_MODEL_ROUTER_KEY}"
1617
deployment_name: "${AZURE_MODEL_ROUTER_DEPLOYMENT}"
1718
parameters:
1819
max_tokens: 1024
1920

2021
baseline:
21-
type: "azure_openai"
22+
type: "openai_compatible"
23+
api_mode: "responses"
2224
endpoint_url: "${AZURE_OPENAI_ENDPOINT}"
2325
api_key: "${AZURE_OPENAI_KEY}"
2426
deployment_name: "${AZURE_BASELINE_DEPLOYMENT}"
@@ -106,7 +108,8 @@ output:
106108
judge:
107109
enabled: true
108110
endpoint:
109-
type: "azure_openai"
111+
type: "openai_compatible"
112+
api_mode: "responses"
110113
endpoint_url: "${AZURE_JUDGE_ENDPOINT}"
111114
api_key: "${AZURE_JUDGE_KEY}"
112115
deployment_name: "${AZURE_JUDGE_DEPLOYMENT}"

‎configs/live_demo.yaml‎

Lines changed: 6 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -8,15 +8,17 @@ evaluation:
88

99
endpoints:
1010
model_router:
11-
type: "azure_openai"
11+
type: "openai_compatible"
12+
api_mode: "chat_completions"
1213
endpoint_url: "${AZURE_MODEL_ROUTER_ENDPOINT}"
1314
api_key: "${AZURE_MODEL_ROUTER_KEY}"
1415
deployment_name: "${AZURE_MODEL_ROUTER_DEPLOYMENT}"
1516
parameters:
1617
max_tokens: 4096
1718

1819
baseline:
19-
type: "azure_openai"
20+
type: "openai_compatible"
21+
api_mode: "responses"
2022
endpoint_url: "${AZURE_OPENAI_ENDPOINT}"
2123
api_key: "${AZURE_OPENAI_KEY}"
2224
deployment_name: "${AZURE_BASELINE_DEPLOYMENT}"
@@ -80,7 +82,8 @@ output:
8082
judge:
8183
enabled: true
8284
endpoint:
83-
type: "azure_openai"
85+
type: "openai_compatible"
86+
api_mode: "responses"
8487
endpoint_url: "${AZURE_JUDGE_ENDPOINT}"
8588
api_key: "${AZURE_JUDGE_KEY}"
8689
deployment_name: "${AZURE_JUDGE_DEPLOYMENT}"

‎docs/faq.md‎

Lines changed: 16 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -37,8 +37,22 @@ cp .env.example .env
3737

3838
| Endpoint | URL format |
3939
|----------|-----------|
40-
| Model Router | `https://<resource>.services.ai.azure.com/models` |
41-
| Azure OpenAI (baseline/judge) | `https://<resource>.openai.azure.com` |
40+
| Model Router | `https://<resource>.services.ai.azure.com/openai/v1` |
41+
| Baseline/judge | `https://<resource>.services.ai.azure.com/openai/v1` |
42+
43+
Use `type: openai_compatible` for these Foundry v1 URLs. Set
44+
`api_mode: chat_completions` for Model Router and `api_mode: responses` for
45+
models whose portal target URI ends in `/responses`. Do not include
46+
`/chat/completions` or `/responses` in the environment variable itself.
47+
48+
### `404 Resource not found`
49+
50+
- Confirm the resource hostname and deployment name exactly match the portal.
51+
- Use the `/openai/v1` base URL, without an operation suffix.
52+
- Ensure `type` is `openai_compatible` and `api_mode` matches the portal target URI.
53+
54+
An invalid hostname usually produces `APIConnectionError`, while a reachable
55+
host with the wrong route or deployment commonly produces `404`.
4256

4357
---
4458

‎docs/how-to-run-live-eval.md‎

Lines changed: 9 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -51,13 +51,13 @@ copy .env.example .env # Windows
5151
Open `.env` and fill in your real values:
5252

5353
```
54-
AZURE_MODEL_ROUTER_ENDPOINT=https://your-resource.services.ai.azure.com/models
54+
AZURE_MODEL_ROUTER_ENDPOINT=https://your-resource.services.ai.azure.com/openai/v1
5555
AZURE_MODEL_ROUTER_KEY=your-model-router-key
5656
AZURE_MODEL_ROUTER_DEPLOYMENT=model-router
57-
AZURE_OPENAI_ENDPOINT=https://your-resource.openai.azure.com
57+
AZURE_OPENAI_ENDPOINT=https://your-resource.services.ai.azure.com/openai/v1
5858
AZURE_OPENAI_KEY=your-azure-openai-key
5959
AZURE_BASELINE_DEPLOYMENT=your-baseline-deployment
60-
AZURE_JUDGE_ENDPOINT=https://your-resource.openai.azure.com
60+
AZURE_JUDGE_ENDPOINT=https://your-resource.services.ai.azure.com/openai/v1
6161
AZURE_JUDGE_KEY=your-azure-openai-key
6262
AZURE_JUDGE_DEPLOYMENT=your-judge-deployment
6363
AZURE_PRICING_REGION=eastus
@@ -71,6 +71,11 @@ AZURE_PRICING_REGION=eastus
7171
- *Azure Portal → your Foundry resource → Keys and Endpoint*
7272
- *Azure Portal → your Azure OpenAI resource → Keys and Endpoint*
7373

74+
Use the resource name exactly as shown by the portal. If the portal gives a
75+
target URI ending in `/chat/completions` or `/responses`, remove that final
76+
operation segment and keep `/openai/v1`. The live presets select the operation
77+
with `api_mode`.
78+
7479
## Step 3: Configure the evaluation (optional)
7580

7681
The default config (`configs/default.yaml`) works out of the box. The settings most people change first:
@@ -80,6 +85,7 @@ The default config (`configs/default.yaml`) works out of the box. The settings m
8085
| Baseline model | `endpoints.baseline.deployment_name` | `gpt-5` | Which model the router is compared against |
8186
| Number of prompts | `evaluation.sample_size` | `null` (all) | How many dataset prompts to use |
8287
| Judge enabled | `judge.enabled` | `true` | Whether to run quality scoring (costs extra API calls) |
88+
| Endpoint API | `api_mode` | Per endpoint | `chat_completions` or `responses` |
8389
| Concurrency | `concurrency.max_parallel_requests` | `5` | How many prompts run in parallel |
8490

8591
To change the baseline model, edit `configs/default.yaml`:

‎pyproject.toml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -10,7 +10,7 @@ readme = "README.md"
1010
license = {text = "MIT"}
1111
requires-python = ">=3.9"
1212
dependencies = [
13-
"openai>=1.0",
13+
"openai>=2.0",
1414
"pyyaml>=6.0",
1515
"matplotlib>=3.7",
1616
"numpy>=1.24",

‎requirements.txt‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,7 @@
66
# in pyproject.toml: `pip install -e .` (core), `.[foundry]`, `.[dev]`, `.[db]`.
77

88
# --- Core (required by scripts/run_eval.py and src/) ---
9-
openai>=1.0
9+
openai>=2.0
1010
pyyaml>=6.0
1111
matplotlib>=3.7
1212
numpy>=1.24

0 commit comments

Comments
 (0)