Repository navigation
Local debugging of azure ml online deployments using arm64-based image #10386
Description
Activity
- addedbugThis issue requires a change to an existing behavior in the product in order to be resolved.This issue requires a change to an existing behavior in the product in order to be resolved.
on Sep 23, 2026 - addedquestionThe issue doesn't require a change to the product in order to be resolved. Most issues start as thatThe issue doesn't require a change to the product in order to be resolved. Most issues start as thatcustomer-reportedIssues that are reported by GitHub users external to the Azure organization.Issues that are reported by GitHub users external to the Azure organization.Auto-AssignAuto assign by botAuto assign by botService AttentionThis issue is responsible by Azure service team.This issue is responsible by Azure service team.
on Sep 23, 2026 Thank you for opening this issue, we will look into it.
x-engineering-agent commented
on Sep 23, 2026 More actionsBug Analysis
Assessment: Sufficient to investigate a local-debugging compatibility bug. The leading source-backed hypothesis is a conda-prefix mismatch, not a demonstrated CLI architecture-selection bug. The reported debugger failure has not been independently reproduced.
Affected extension:
ml2.45.0, maintained undersrc/machinelearningservices/(azext_mlv2). Keep this work on #10386; no tracker issue is needed. Target branch:main.Reported reproduction and behavior
az ml online-deployment create --file endpoints/online/managed/sample/blue-deployment.yml --local --vscode-debugThe author also supplied this command with
--debug. Context: Azure CLI/core 2.90.0,ml2.45.0, Apple M5 Pro/ARM64. The standard sample image produces an AMD64 local environment with unusable LLM performance. Replacing the image withcondaforge/miniforge3:latestproduces the desired ARM64 environment, but VS Code cannot attach the debugger. Expected: native ARM64 local execution with working VS Code debugging.The supplied VS Code output starts the debug adapter with
/opt/conda/bin/pythonand reports the active scoring interpreter as/opt/conda/envs/inf-conda-env/bin/python. It contains a Mamba-manager warning but no attach exception. No author follow-up adds further diagnostic evidence. The image digest and exact attach failure remain unconfirmed; do not treat an informational adapter-start log or the Mamba warning as a proven root cause.Current source evidence
- Extension metadata confirms version 2.45.0; requirements pins
azure-ai-ml==1.35.0. ml_online_deployment_createforwardslocal,local_enable_gpu, andvscode_debugintoMLClient.begin_create_or_update. The examined CLI wrapper does not force AMD64 or implement debugger attachment.- In the pinned SDK,
LocalEndpointConstantshardcodes the inference executable directory and Python path under/opt/miniconda/envs/inf-conda-env, unlike the author's/opt/condalayout. AzureMlImageContextassigns that hardcoded directory toAZUREML_INFERENCE_PYTHON_PATH. Separately,Settings.to_dictassigns the hardcoded executable topython.defaultInterpreterPathand creates the Azure ML local-inference attach configuration. Selecting a different VS Code interpreter alone does not change the container environment variable.DockerfileResolver._constructkeeps the selected base image and installs debugpy into the named conda environment.VSCodeClientpasses the environment through to the devcontainer configuration. Thus successful ARM64 image creation does not establish that the inference/debug launch uses a valid interpreter.
Implementation handoff and acceptance criteria
- Confirm whether the inference/debug launch actually consumes the invalid
/opt/minicondapath with an/opt/condaimage before choosing a fix. Inspect the generated devcontainer configuration, inference environment and Azure ML debug-launch integration; do not assume ARM64 itself or debugpy is broken. - Resolve both inference and debugger interpreter paths from the selected image's actual conda environment, through the authoritative owning source. Do not globally replace one hardcoded prefix with another, infer a container path from host architecture, force AMD64 emulation, or silently change the user's base image. Preserve the existing
/opt/minicondalayout and non-debug/managed-deployment behavior. - Source ownership is a gate: the concrete path construction above belongs to the
azure-ai-mlSDK, not an AAZ command model. The extension files inspected also carry AutoRest-generated headers, including files belowmanual/. Locate and change their durable source/generation inputs and use the owning generation/release workflow; never patch generated output, installed packages, or vendor SDK internals. If the necessary SDK/source change is outside the durable job's approved scope, report that precise upstream prerequisite/blocker instead of manufacturing a same-repository fix. Do not create another issue or dispatch another backend. - Add focused regression coverage in the owning source for
/opt/condaand/opt/minicondalayouts, consistent inference/debugger paths, unchanged image selection, and unchanged forwarding of local/debug flags. Preserve the Azure ML attach protocol. Use the repository's existing validation only in the authorized implementation environment; no infrastructure provisioning is requested by this triage. - For a downstream extension release, consume the verified source fix through its normal dependency/generation process and maintain version/history/compatibility metadata. Do not manually edit
src/index.json. Do not claim working ARM64 debugger attachment without evidence.
This is a source-inspection handoff, not a claim of a completed fix or executed tests. The implementation backend is the durable Foundry job only.
Mandatory Codegen execution protocol
Before editing implementation files, determine whether the affected
machinelearningservicescommand is AAZ-generated. Files underaaz/<profile>/are generated output and must never be patched directly, including by an AI agent. Check outAzure/aazbesideAzure/azure-rest-api-specs,Azure/aaz-dev-tools, and the downstream repository. API-schema defects start in the specification; command naming, grouping, arguments, API-version selection, help, and examples belong in the durableAzure/aazcommand model; non-modelable client behavior belongs in a handwritten subclass or wrapper incustom.py, registered fromcommands.py. X Engineering Agent creates and promotes the corresponding durableAzure/aazsource pull request before it promotes downstream generated output.Follow the Azure CLI repository's Codegen workflow and the aaz-dev setup documentation. Set up the checked-out repositories with
azdev setup. Usegenerateonly when importing or redesigning command models from Swagger/TypeSpec. For an existing module whose durableAzure/aazmodel has been updated, render that model withregenerate:aaz-dev cli regenerate --name machinelearningservices --cli-extension-path <azure-cli-extensions> # New/imported command model only: aaz-dev cli generate --spec <specification-name> --module machinelearningservices
You MUST actually run the generator; do not merely describe it or imitate its output. If the AAZ/specification checkout, local source change, credentials, or generator is unavailable, stop and report the blocker instead of editing generated files. Inspect
_aaz_infoprovenance and the complete regenerated diff, then run focusedazdev style,azdev linter, andazdev testvalidation. For an extension, also update its version andHISTORY.rst, preserveazext_metadata.jsoncompatibility, and let release automation updatesrc/index.json.PR title & description format (required)
This repo enforces a PR format (guide). Please author the PR exactly as follows or CI's Check the Format of Pull Request Title and Content will fail.
Use this EXACT PR title (copy verbatim, do not reword):
[Machinelearningservices] Fix #10386: `az ml online-deployment create`: Resolve local debugging conda paths for ARM64 imagesKeep the backticks around the command and the
Fix #10386:prefix. You may only adjust the wording after the command (the final summary) if the fix changes; the[Machinelearningservices]prefix, issue link, and backticked command must stay.Description — follow the PR template and fill in:
- Link the issue — start the Description with a closing keyword so the PR auto-links and closes it:
Fixes #10386. - Related command — the
az ...command this affects. - Description (mandatory) — why the bug happens, what you changed, and the resulting behavior.
- Testing Guide — example command(s) showing the fix works.
- History Notes — leave the title to drive the history note, or add extra lines in the same format (component in brackets + the command in backticks), e.g.
[Machinelearningservices] `az <command>`: <note>. - Keep the template checklist and tick the items you've satisfied.
- Extension metadata confirms version 2.45.0; requirements pins
- removedquestionThe issue doesn't require a change to the product in order to be resolved. Most issues start as thatThe issue doesn't require a change to the product in order to be resolved. Most issues start as that
on Sep 23, 2026 Bug Analysis
Assessment: Sufficient to investigate a local-debugging compatibility bug. The leading source-backed hypothesis is a conda-prefix mismatch, not a demonstrated CLI architecture-selection bug. The reported debugger failure has not been independently reproduced.
Affected extension:
ml2.45.0, maintained undersrc/machinelearningservices/(azext_mlv2). Keep this work on #10386; no tracker issue is needed. Target branch:main.Reported reproduction and behavior
az ml online-deployment create --file endpoints/online/managed/sample/blue-deployment.yml --local --vscode-debugThe author also supplied this command with
--debug. Context: Azure CLI/core 2.90.0,ml2.45.0, Apple M5 Pro/ARM64. The standard sample image produces an AMD64 local environment with unusable LLM performance. Replacing the image withcondaforge/miniforge3:latestproduces the desired ARM64 environment, but VS Code cannot attach the debugger. Expected: native ARM64 local execution with working VS Code debugging.The supplied VS Code output starts the debug adapter with
/opt/conda/bin/pythonand reports the active scoring interpreter as/opt/conda/envs/inf-conda-env/bin/python. It contains a Mamba-manager warning but no attach exception. No author follow-up adds further diagnostic evidence. The image digest and exact attach failure remain unconfirmed; do not treat an informational adapter-start log or the Mamba warning as a proven root cause.Current source evidence
* [Extension metadata](https://github.com/Azure/azure-cli-extensions/blob/b601ece641dcd01698bab48ec690956f0fd4f060/src/machinelearningservices/setup.py#L12-L14) confirms version 2.45.0; [requirements](https://github.com/Azure/azure-cli-extensions/blob/b601ece641dcd01698bab48ec690956f0fd4f060/src/machinelearningservices/azext_mlv2/manual/requirements.txt) pins `azure-ai-ml==1.35.0`. * [`ml_online_deployment_create`](https://github.com/Azure/azure-cli-extensions/blob/b601ece641dcd01698bab48ec690956f0fd4f060/src/machinelearningservices/azext_mlv2/manual/custom/online_deployment.py#L133-L140) forwards `local`, `local_enable_gpu`, and `vscode_debug` into `MLClient.begin_create_or_update`. The examined CLI wrapper does not force AMD64 or implement debugger attachment. * In the pinned SDK, [`LocalEndpointConstants`](https://github.com/Azure/azure-sdk-for-python/blob/59fb7a9d52b4c85ac5076d31759f5c3a501072e4/sdk/ml/azure-ai-ml/azure/ai/ml/constants/_endpoint.py#L73-L80) hardcodes the inference executable directory and Python path under `/opt/miniconda/envs/inf-conda-env`, unlike the author's `/opt/conda` layout. * [`AzureMlImageContext`](https://github.com/Azure/azure-sdk-for-python/blob/59fb7a9d52b4c85ac5076d31759f5c3a501072e4/sdk/ml/azure-ai-ml/azure/ai/ml/_local_endpoints/azureml_image_context.py#L68-L72) assigns that hardcoded directory to `AZUREML_INFERENCE_PYTHON_PATH`. Separately, [`Settings.to_dict`](https://github.com/Azure/azure-sdk-for-python/blob/59fb7a9d52b4c85ac5076d31759f5c3a501072e4/sdk/ml/azure-ai-ml/azure/ai/ml/_local_endpoints/vscode_debug/devcontainer_properties.py#L173-L196) assigns the hardcoded executable to `python.defaultInterpreterPath` and creates the Azure ML local-inference attach configuration. Selecting a different VS Code interpreter alone does not change the container environment variable. * [`DockerfileResolver._construct`](https://github.com/Azure/azure-sdk-for-python/blob/59fb7a9d52b4c85ac5076d31759f5c3a501072e4/sdk/ml/azure-ai-ml/azure/ai/ml/_local_endpoints/dockerfile_resolver.py#L88-L120) keeps the selected base image and installs debugpy into the named conda environment. [`VSCodeClient`](https://github.com/Azure/azure-sdk-for-python/blob/59fb7a9d52b4c85ac5076d31759f5c3a501072e4/sdk/ml/azure-ai-ml/azure/ai/ml/_local_endpoints/vscode_debug/vscode_client.py#L14-L32) passes the environment through to the devcontainer configuration. Thus successful ARM64 image creation does not establish that the inference/debug launch uses a valid interpreter.Implementation handoff and acceptance criteria
1. Confirm whether the inference/debug launch actually consumes the invalid `/opt/miniconda` path with an `/opt/conda` image before choosing a fix. Inspect the generated devcontainer configuration, inference environment and Azure ML debug-launch integration; do not assume ARM64 itself or debugpy is broken. 2. Resolve both inference and debugger interpreter paths from the selected image's actual conda environment, through the authoritative owning source. Do not globally replace one hardcoded prefix with another, infer a container path from host architecture, force AMD64 emulation, or silently change the user's base image. Preserve the existing `/opt/miniconda` layout and non-debug/managed-deployment behavior. 3. **Source ownership is a gate:** the concrete path construction above belongs to the `azure-ai-ml` SDK, not an AAZ command model. The extension files inspected also carry AutoRest-generated headers, including files below `manual/`. Locate and change their durable source/generation inputs and use the owning generation/release workflow; never patch generated output, installed packages, or vendor SDK internals. If the necessary SDK/source change is outside the durable job's approved scope, report that precise upstream prerequisite/blocker instead of manufacturing a same-repository fix. Do not create another issue or dispatch another backend. 4. Add focused regression coverage in the owning source for `/opt/conda` and `/opt/miniconda` layouts, consistent inference/debugger paths, unchanged image selection, and unchanged forwarding of local/debug flags. Preserve the Azure ML attach protocol. Use the repository's existing validation only in the authorized implementation environment; no infrastructure provisioning is requested by this triage. 5. For a downstream extension release, consume the verified source fix through its normal dependency/generation process and maintain version/history/compatibility metadata. Do not manually edit `src/index.json`. Do not claim working ARM64 debugger attachment without evidence.This is a source-inspection handoff, not a claim of a completed fix or executed tests. The implementation backend is the durable Foundry job only.
Mandatory Codegen execution protocol
Before editing implementation files, determine whether the affected
machinelearningservicescommand is AAZ-generated. Files underaaz/<profile>/are generated output and must never be patched directly, including by an AI agent. Check outAzure/aazbesideAzure/azure-rest-api-specs,Azure/aaz-dev-tools, and the downstream repository. API-schema defects start in the specification; command naming, grouping, arguments, API-version selection, help, and examples belong in the durableAzure/aazcommand model; non-modelable client behavior belongs in a handwritten subclass or wrapper incustom.py, registered fromcommands.py. X Engineering Agent creates and promotes the corresponding durableAzure/aazsource pull request before it promotes downstream generated output.Follow the Azure CLI repository's Codegen workflow and the aaz-dev setup documentation. Set up the checked-out repositories with
azdev setup. Usegenerateonly when importing or redesigning command models from Swagger/TypeSpec. For an existing module whose durableAzure/aazmodel has been updated, render that model withregenerate:aaz-dev cli regenerate --name machinelearningservices --cli-extension-path
New/imported command model only:
aaz-dev cli generate --spec --module machinelearningservices
You MUST actually run the generator; do not merely describe it or imitate its output. If the AAZ/specification checkout, local source change, credentials, or generator is unavailable, stop and report the blocker instead of editing generated files. Inspect
_aaz_infoprovenance and the complete regenerated diff, then run focusedazdev style,azdev linter, andazdev testvalidation. For an extension, also update its version andHISTORY.rst, preserveazext_metadata.jsoncompatibility, and let release automation updatesrc/index.json.PR title & description format (required)
This repo enforces a PR format (guide). Please author the PR exactly as follows or CI's Check the Format of Pull Request Title and Content will fail.
Use this EXACT PR title (copy verbatim, do not reword):
[Machinelearningservices] Fix #10386: `az ml online-deployment create`: Resolve local debugging conda paths for ARM64 imagesKeep the backticks around the command and the
Fix #10386:prefix. You may only adjust the wording after the command (the final summary) if the fix changes; the[Machinelearningservices]prefix, issue link, and backticked command must stay.Description — follow the PR template and fill in:
* **Link the issue** — start the Description with a closing keyword so the PR auto-links and closes it: `Fixes #10386`. * **Related command** — the `az ...` command this affects. * **Description** _(mandatory)_ — why the bug happens, what you changed, and the resulting behavior. * **Testing Guide** — example command(s) showing the fix works. * **History Notes** — leave the title to drive the history note, or add extra lines in the same format (component in brackets + the command in backticks), e.g. `` [Machinelearningservices] `az <command>`: <note> ``. * Keep the template checklist and tick the items you've satisfied.Providing arm64-compatible base images for the ML service could be the most straight forward solution. This would avoid having to force users to use other image sources.
- addedRequest X Engineering AgentRequest X Engineering Agent testing and reviewRequest X Engineering Agent testing and review
on Oct 2, 2026 microsoft-github-policy-service commented
on Oct 2, 2026 ContributorMore actionsThanks for the feedback! We are routing this to the appropriate team for follow-up. cc azureml-github, @Azure/act-experience-enablement-squad.
x-engineering-agent commented
on Oct 2, 2026 More actionsOwnership follow-up
Official ARM64 Azure ML base images would address the image-availability concern without requiring users to choose another image provider. That request needs the Azure ML image-publishing owners; it is separate from the local debugger compatibility issue.
The extension is now ml 2.45.1, and its current dependency pin is
azure-ai-ml==1.35.1. That SDK still hardcodes the inference environment under/opt/miniconda, and its VS Code configuration selects that interpreter. This keeps the alternate-Conda-prefix hypothesis relevant, but does not establish the cause of the reported attach failure.The implementation prerequisite belongs to the
azure-ai-mlSDK in Azure/azure-sdk-for-python; official image availability belongs to the Azure ML image owners. I am not requeuing an extensions-only implementation or replacing the selected image, because neither would resolve those ownership prerequisites. A downstream extension update should consume a verified SDK fix through the normal release process, preserving existing Miniconda behavior and user-selected images.Maintainers: please coordinate the SDK compatibility investigation and the official ARM64-image request with those owners. This issue remains open; the current Agent analysis request is complete. Automation can be requested again when an actionable extension change or an upstream release is available.
Posted by x-engineering-agent (Fixer)
- removedRequest X Engineering AgentRequest X Engineering Agent testing and reviewRequest X Engineering Agent testing and review
on Oct 2, 2026
Describe the bug
Expected Behavior
I expect to be able to follow https://learn.microsoft.com/en-us/azure/machine-learning/how-to-debug-managed-online-endpoints-visual-studio-code?view=azureml-api-2&tabs=cli and get an arm64-based local debugging environment.
Actual Behavior
I get a amd64-based local debugging environment, which on apple silicon is unusable performance-wise
Steps to Reproduce the Problem
condaforge/miniforge3:latestand get an arm64-based local debugging environment. Fast enough for LLMs, but the debugger is unable to attach. Debug logs of the creation process attached. VSCode logs attached in errors section (but no errors visible)Specifications
Related command
Errors
Issue script & Debug output
Expected behavior
I expect to be able to follow https://learn.microsoft.com/en-us/azure/machine-learning/how-to-debug-managed-online-endpoints-visual-studio-code?view=azureml-api-2&tabs=cli and get an arm64-based local debugging environment.
Environment Summary
Additional context
No response