Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
67 commits
Select commit Hold shift + click to select a range
26c37c8
feat(litert): run the LiteRT backend on iOS
anhappdev Aug 27, 2026
874dddf
build(litert): build and embed the LiteRT framework in the iOS app
anhappdev Aug 27, 2026
3fffeae
ci(litert): cover the LiteRT backend in the iOS matrix
anhappdev Aug 27, 2026
1298870
fix(litert): propagate shapes before creating tensor buffers
anhappdev Aug 27, 2026
bfd93a5
feat(litert): default the iOS vision benchmarks to Metal
anhappdev Aug 27, 2026
826241e
docs(litert): record the iOS delegate defaults and the LLM memory limit
anhappdev Aug 27, 2026
7663912
ci: raise the iOS device step timeout so the job cap can apply
anhappdev Aug 27, 2026
2761e3b
docs(litert): correct the iOS LLM memory limit
anhappdev Aug 27, 2026
2218a2d
Merge branch 'litert' into litert-ios
anhappdev Aug 28, 2026
7ed1453
feat(litert): claim stable diffusion and fit the 1B LLMs on iOS
anhappdev Aug 28, 2026
cc11c47
style(litert): clang-format the prefill signature call site
anhappdev Aug 28, 2026
9b1b7fc
fix(litert): keep stable diffusion inside the iOS memory limit
anhappdev Aug 28, 2026
4cfd7a4
ci: stop the device log download failing silently
anhappdev Aug 28, 2026
be88e18
fix(litert): guard ReleaseModel so Android has no unused function
anhappdev Aug 28, 2026
ca8f0a3
feat(litert): log the remaining iOS memory budget around each phase
anhappdev Aug 28, 2026
19eda8f
fix(litert): hand freed pages back to the OS before the SD decode
anhappdev Aug 28, 2026
98b3e2f
feat(litert): send the memory numbers to os_log so CI can diagnose it…
anhappdev Aug 28, 2026
1b2a49c
fix(litert): keep only one stable diffusion phase resident on Apple
anhappdev Aug 28, 2026
3851324
feat(litert): instrument the LLM pipeline's memory and bucket choice
anhappdev Aug 28, 2026
71d8818
fix(litert): stop offering Metal for the LLM benchmarks on iOS
anhappdev Aug 28, 2026
a964637
feat(litert): offer the llm-* benchmarks only where they fit
anhappdev Aug 28, 2026
fe4f511
build: raise every iOS framework minimum to 14.0
anhappdev Aug 28, 2026
b923b03
Revert "feat(litert): offer the llm-* benchmarks only where they fit"
anhappdev Aug 28, 2026
897f89c
ci(ios): reproduce Xcode Cloud's App Store Connect check in GitHub CI
anhappdev Aug 31, 2026
cb2c54a
Revert "ci(ios): reproduce Xcode Cloud's App Store Connect check in G…
anhappdev Aug 31, 2026
4922d87
fix(litert): ship the Metal accelerator as a framework, not a bare dylib
anhappdev Aug 31, 2026
ebd667f
Merge branch 'litert' into litert-ios
anhappdev Sep 8, 2026
39b3f4a
feat(ios): request the increased memory limit and extended virtual ad…
anhappdev Sep 8, 2026
842beb1
feat(litert): offer Metal for llm-1b and stable diffusion on iOS
anhappdev Sep 8, 2026
fd6b13c
docs(litert): record what offering the iOS Metal choices costs
anhappdev Sep 8, 2026
f331cac
fix(litert): make the backend build and export symbols on macOS
anhappdev Sep 8, 2026
79985e3
docs(litert): replace the iOS Metal guesses with host measurements
anhappdev Sep 8, 2026
42cc3f3
docs(litert): name what blocks stable diffusion on Metal — the batch dim
anhappdev Sep 8, 2026
aeb5f47
docs(litert): correct why stable diffusion cannot use Metal
anhappdev Sep 8, 2026
7858455
docs(litert): add the CPU accuracy control for the Metal LLM run
anhappdev Sep 8, 2026
5d63686
feat(litert): select Metal for the llm-* benchmarks on iOS
anhappdev Sep 8, 2026
3f3baf9
fix(litert): actually release hardware tensor buffers on teardown
anhappdev Sep 8, 2026
764b34a
fix(tasks): give the llm-* quick runs time to meet min_query_count
anhappdev Sep 8, 2026
000a915
Merge branch 'litert' of https://github.com/mlcommons/mobile_app_open…
anhappdev Sep 8, 2026
1cb8840
style: satisfy the clang-format and markdownlint checks
anhappdev Sep 8, 2026
ad0e1b8
feat(litert): make the stable diffusion models run on the Metal delegate
anhappdev Sep 9, 2026
482fb5a
fix(litert): make the SD converter's equivalence check sound
anhappdev Sep 9, 2026
e5817b8
docs(litert): record what moving the LiteRT pin to 2.2.0 would cost
anhappdev Sep 9, 2026
9648e92
docs(litert): the SD rewrite buys nothing on CPU at the pinned 2.1.5
anhappdev Sep 9, 2026
a492e8a
feat: move LiteRT to 2.2.0 so stable diffusion can run on Metal
anhappdev Sep 9, 2026
e5438ad
chore(litert): run the SD converter under uv with pinned dependencies
anhappdev Sep 9, 2026
14f4350
Merge branch 'litert-ios' into litert-ios-2.2.0
anhappdev Sep 9, 2026
2986839
feat(litert): run stable_diffusion on Metal
anhappdev Sep 9, 2026
f52bc0a
fix(litert): version-stamp the accelerator cache, and correct why SD …
anhappdev Sep 9, 2026
6f9918e
docs(litert): correct claims the review found overstated
anhappdev Sep 9, 2026
b2a01e4
fix(build): repair the Android and Windows builds after the 2.2.0 pin…
anhappdev Sep 9, 2026
a7603ac
fix(build): turn off the layering check for host-tool compiles
anhappdev Sep 9, 2026
5cae730
fix(build): patch LiteRT's own Logistic, and declare tsl's status_macros
anhappdev Sep 9, 2026
2bb48b9
fix(build): finish repairing the Android build after the 2.2.0 pin move
anhappdev Sep 9, 2026
a9298cb
fix(litert): let stable_diffusion fit on Metal by re-quantising FC we…
anhappdev Sep 9, 2026
c2f5de7
fix(litert): point the Metal SD choice at the per-axis exports
anhappdev Sep 9, 2026
3ff6c3a
style(litert): satisfy clang-format in the SD phase split
anhappdev Sep 10, 2026
589296e
ci(ios): raise the device step timeout to 120 minutes
anhappdev Sep 10, 2026
13858af
fix(litert): name the _v2 SD models in the filename settings too
anhappdev Sep 10, 2026
3872285
fix(litert): keep the SD models resident on Metal, drop the phase reb…
anhappdev Sep 10, 2026
b2f9f16
style(litert): reformat the GPU residency guard with clang-format 16
anhappdev Sep 10, 2026
c1d70b3
fix(ios): declare the accelerator framework's real minimum OS
anhappdev Sep 11, 2026
3a0ec36
docs(litert): correct the README where SD on Metal overtook it
anhappdev Sep 11, 2026
46686c4
Merge remote-tracking branch 'github-https/litert' into litert-ios
anhappdev Sep 22, 2026
056213c
build(apple): give the CoreML Swift library an explicit language mode
anhappdev Sep 22, 2026
f4e8400
fix(ios): adopt the UIScene lifecycle so iOS 27 does not trap at launch
anhappdev Sep 22, 2026
f3d44f2
fix(ifeval): stop reading before the start of a response
anhappdev Sep 22, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions .bazelrc
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,23 @@ build --per_file_copt=mobile_back_tflite/cpp/backend_tflite/llm_pipeline.cc@-fpe
# tools resolved at fetch time still work.
build --incompatible_strict_action_env

# Abseil does not survive clang's layering check when it is built for the exec
# configuration, i.e. as part of a host tool such as protoc:
#
# absl/status/statusor.h:58:10: error: module com_google_absl//absl/status:
# statusor does not depend on a module exporting 'absl/types/span.h'
#
# The dependency is declared -- //absl/types:span is in statusor's deps and its
# .cppmap is on the command line -- so this is the runner's clang disagreeing
# with Abseil's own module layering, not a missing edge we could add. It is a
# hygiene diagnostic over third-party code that produces no artifact, so it is
# turned off rather than worked around.
#
# It went unnoticed until the LiteRT 2.2.0 pin move: the action was a remote
# cache hit on every previous green run, and only the cold cache the new pins
# forced actually compiled it. Target configurations keep the check.
build --host_features=-layering_check

# This flag is required by tensorflow
common --experimental_repo_remote_exec

Expand Down
24 changes: 22 additions & 2 deletions .github/workflows/ios-build-test-macos.yml
Original file line number Diff line number Diff line change
Expand Up @@ -41,9 +41,15 @@ jobs:
- backend: "tflite"
with_apple: 0
with_tflite: 1
with_litert: 0
- backend: "apple"
with_apple: 1
with_tflite: 0
with_litert: 0
- backend: "litert"
with_apple: 0
with_tflite: 0
with_litert: 1
env:
PERF_TEST: true
BAZEL_OUTPUT_ROOT_ARG: "--output_user_root=/tmp/bazel_output"
Expand All @@ -55,6 +61,7 @@ jobs:
FLUTTER_IOS_TEST_PACKAGE: ios_tests-${{ matrix.backend }}-${{ needs.computed.outputs.build_number }}.zip
WITH_APPLE: ${{ matrix.with_apple }}
WITH_TFLITE: ${{ matrix.with_tflite }}
WITH_LITERT: ${{ matrix.with_litert }}
WITH_PIXEL: 0
WITH_MEDIATEK: 0
WITH_QTI: 0
Expand Down Expand Up @@ -229,6 +236,9 @@ jobs:
runs-on: ubuntu-22.04
permissions:
contents: read
# The litert row runs the llm-* benchmarks, whose go-button check validates
# every delegate choice's model file (two ~1.2 GB LLM exports), so it needs
# the same headroom as the Android unified job.
timeout-minutes: 135
# BrowserStack allows only 2 parallel sessions on this account, and a
# superseded run's device tests starve the current one -- this job has
Expand All @@ -254,6 +264,8 @@ jobs:
device: "iPhone 16 Pro-18"
- backend: "apple"
device: "iPhone 16 Pro-18"
- backend: "litert"
device: "iPhone 16 Pro-18"
env:
TEST_PACKAGE_NAME: ios_tests-${{ matrix.backend }}-${{ needs.computed.outputs.build_number }}.zip
BROWSERSTACK_LOGS_DIR: /tmp/browserstack-device-logs
Expand Down Expand Up @@ -282,11 +294,19 @@ jobs:
BROWSERSTACK_DEVICES: >-
["${{ matrix.device }}"]
with:
timeout_minutes: 60
# This is the timeout that actually fires: the job-level one never
# gets a chance, because the step waits out retry_wait_seconds even
# when it will not retry (retry_on_exit_code only covers a trigger
# rejection). Keep it below the job cap minus retry_wait_seconds.
# BrowserStack reports a build as "running" while it is still queued
# for a device, so this budget covers queue time as well as the run.
timeout_minutes: 80
# Exit 9 is a BrowserStack trigger failure, which is what
# BROWSERSTACK_ALL_PARALLELS_IN_USE produces when the account's
# session budget is busy - retrying it just waits for a free slot.
# Every other failure still fails fast.
# Every other failure still fails fast. A rejection returns
# immediately, but each retry still costs retry_wait_seconds, so a
# long queue can burn into the job cap above.
max_attempts: 8
retry_wait_seconds: 300
retry_on_exit_code: 9
Expand Down
23 changes: 17 additions & 6 deletions .github/workflows/scripts/browserstack-app-automate.sh
Original file line number Diff line number Diff line change
Expand Up @@ -125,16 +125,27 @@
if [[ -n "$test_id" && "$test_id" != "null" && -n "$device_log_url" && "$device_log_url" != "null" ]]; then
echo "Found last test case $test_id with device log URL"

# Download device logs using the extracted URL
# Download device logs using the extracted URL. -f so an HTTP error is a
# failure instead of an error page written to the log, and retries because
# the log is not always served the instant the session ends.
local log_file="$LOGS_DIR/${test_id}.log"
echo "Downloading device log to $log_file"
curl -s -u "$CREDENTIALS" -X GET "$device_log_url" -o "$log_file"

if [ -f "$log_file" ]; then
echo "Device logs downloaded successfully to $log_file"
if curl -sf --retry 3 --retry-delay 5 -u "$CREDENTIALS" -X GET \
"$device_log_url" -o "$log_file" && [ -s "$log_file" ]; then

Check failure on line 134 in .github/workflows/scripts/browserstack-app-automate.sh

View check run for this annotation

SonarQubeCloud / SonarCloud Code Analysis

Use '[[' instead of '[' for conditional tests. The '[[' construct is safer and more feature-rich.

See more on https://sonarcloud.io/project/issues?id=mlcommons_mobile_app_open&issues=AaBGY5wB4YzCRoHRCTut&open=AaBGY5wB4YzCRoHRCTut&pullRequest=1175
echo "Device logs downloaded successfully to $log_file ($(wc -c <"$log_file") bytes)"
else
echo "Failed to download device logs for test case $test_id"
# Test -s, not -f: a zero-byte log is what this used to leave behind, and
# it reads as "the run produced no output" rather than "the download
# failed". Leave a note in its place so the artifact explains itself --
# if-no-files-found on the upload step only catches a missing file.
echo "Failed to download device logs for test case $test_id" \
| tee "$log_file"
echo " session: $session_id" >> "$log_file"
echo " url: $device_log_url" >> "$log_file"
fi
else
echo "No device log URL in the session response for build $build_id" \
| tee "$LOGS_DIR/no-device-log-${session_id}.txt"
fi
}

Expand Down
43 changes: 32 additions & 11 deletions WORKSPACE
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ http_archive(
# rules_python that XLA now brings in: rules_apple's plisttool is generated
# from a bootstrap template containing %interpreter_args%, which the older
# rules leave unsubstituted, so it fails to parse as Python. These are the
# same versions LiteRT 2.1.5 pins for this dependency set.
# same versions LiteRT 2.2.0 pins for this dependency set.
http_archive(
name = "build_bazel_rules_apple",
sha256 = "a78f26c22ac8d6e3f3fcaad50eace4d9c767688bd7254b75bdf4a6735b299f6a",
Expand All @@ -51,6 +51,23 @@ http_archive(
url = "https://github.com/bazelbuild/apple_support/releases/download/1.23.1/apple_support.1.23.1.tar.gz",
)

# The rules_cc rules_python's py_repositories() would bring in anyway, patched.
# Declared here so it is fetched with the patch rather than without it; the
# declaration below it is a maybe(), so ours wins.
#
# Moving TensorFlow to LiteRT 2.2.0's commit pulled this chain -- XLA, then
# rules_ml_toolchain, then rules_python -- forward to a rules_cc whose
# use_cc_toolchain() is mandatory. That breaks the Android build during
# analysis. See patches/rules_cc_optional_cc_toolchain.patch for the detail.
http_archive(
name = "rules_cc",
patch_args = ["-p1"],
patches = ["//patches:rules_cc_optional_cc_toolchain.patch"],
sha256 = "b8b918a85f9144c01f6cfe0f45e4f2838c7413961a8ff23bc0c6cdf8bb07a3b6",
strip_prefix = "rules_cc-0.1.5",
url = "https://github.com/bazelbuild/rules_cc/releases/download/0.1.5/rules_cc-0.1.5.tar.gz",
)

http_archive(
name = "bazel_features",
sha256 = "c26b4e69cf02fea24511a108d158188b9d8174426311aac59ce803a78d107648",
Expand Down Expand Up @@ -131,21 +148,21 @@ http_archive(
patch_args = ["-p1"],
patches = [
# Add patches for adding png in tflite evaluation code
# Channel detection and the unsigned-char decode buffer, formerly
# png-with-number-of-channels-detected.patch and use_unsigned_char.patch,
# are folded into this one.
"//:flutter/third_party/enable-png-in-tensorflow-lite-tools-evaluation.patch",
"//:flutter/third_party/png-with-number-of-channels-detected.patch",
"//:flutter/third_party/use_unsigned_char.patch",
# Fix tensorflow not being able to read image files on Windows
"//:flutter/third_party/tensorflow-fix-file-opening-mode-for-Windows.patch",
"//:patches/litert_logistic_fp16_msvc.patch",
"//:patches/litert-internal-visibility.diff",
# Fix for LiteRT crashing on close when using OpenCL accelerator
"//:patches/custom_buffer_teardown.patch",
# CoreML delegate calls RepeatedField::resize, which does not exist
"//:patches/litert_coreml_repeatedfield_resize.patch",
],
sha256 = "7d0313c4851deb18af6f5f2dbc002bf01293583b87b819b0949ee33dcfe2d91b",
strip_prefix = "LiteRT-2.1.5",
sha256 = "6d2ce16738199adc5a3cdde76c3c6a6dac636d3b52a1d7790ea524fb0d59f7fc",
strip_prefix = "LiteRT-2.2.0",
urls = [
"https://github.com/google-ai-edge/LiteRT/archive/v2.1.5.tar.gz",
"https://github.com/google-ai-edge/LiteRT/archive/v2.2.0.tar.gz",
],
)

Expand All @@ -161,13 +178,17 @@ tensorflow_source_repo(
patches = [
"//:flutter/third_party/tf-eigen.patch",
"//patches:tf_coreml_repeatedfield_resize.patch",
"//patches:tf_logistic_fp16_msvc.patch",
"//patches:tf_nnapi_no_mmap_sharing.patch",
"//patches:tf_portable_no_onednn_env_vars.patch",
] + PATCH_FILE,
sha256 = "879cf25692d50c60315a4dd3929dccd923d4c44a2c4b95ebb483666d2c16a22a",
strip_prefix = "tensorflow-6d40c20cdfe385746c31da6227b95722f5ece342",
# The commit LiteRT 2.2.0 pins. LiteRT's GPU delegate needs
# @com_google_absl//absl/status:status_macros, which only exists in the
# Abseil this TensorFlow's XLA brings in, so the two move together.
sha256 = "c3c552414ab2e59e72511a21c1df566346a7c8f160909325edec6d1ff403d69d",
strip_prefix = "tensorflow-bcdab1a62e138c8f8784a7477c0be8af6dd0bd0a",
urls = [
"https://github.com/tensorflow/tensorflow/archive/6d40c20cdfe385746c31da6227b95722f5ece342.tar.gz",
"https://github.com/tensorflow/tensorflow/archive/bcdab1a62e138c8f8784a7477c0be8af6dd0bd0a.tar.gz",
],
)

Expand Down
48 changes: 42 additions & 6 deletions flutter/assets/tasks.pbtxt
Original file line number Diff line number Diff line change
Expand Up @@ -417,7 +417,13 @@ task {
quick {
min_query_count: 10
min_duration: 10
max_duration: 40
# 40s could never satisfy min_query_count on a 1B model: measured on an
# iPhone 16 Pro, llm-1b runs about 10.1s per query on the Metal delegate
# (24.3 tok/s), so ten queries need ~101s and the run ended invalid --
# isMinQueryMet false -- on both the CPU and the GPU delegate. A run
# stops as soon as min_duration AND min_query_count are both met, so
# this ceiling only costs time on devices that were failing anyway.
max_duration: 150
}
rapid {
min_query_count: 6
Expand Down Expand Up @@ -477,7 +483,13 @@ task {
quick {
min_query_count: 10
min_duration: 10
max_duration: 40
# 40s could never satisfy min_query_count on a 1B model: measured on an
# iPhone 16 Pro, llm-1b runs about 10.1s per query on the Metal delegate
# (24.3 tok/s), so ten queries need ~101s and the run ended invalid --
# isMinQueryMet false -- on both the CPU and the GPU delegate. A run
# stops as soon as min_duration AND min_query_count are both met, so
# this ceiling only costs time on devices that were failing anyway.
max_duration: 150
}
rapid {
min_query_count: 6
Expand Down Expand Up @@ -657,7 +669,13 @@ task {
quick {
min_query_count: 10
min_duration: 10
max_duration: 40
# 40s could never satisfy min_query_count on a 1B model: measured on an
# iPhone 16 Pro, llm-1b runs about 10.1s per query on the Metal delegate
# (24.3 tok/s), so ten queries need ~101s and the run ended invalid --
# isMinQueryMet false -- on both the CPU and the GPU delegate. A run
# stops as soon as min_duration AND min_query_count are both met, so
# this ceiling only costs time on devices that were failing anyway.
max_duration: 150
}
rapid {
min_query_count: 6
Expand Down Expand Up @@ -717,7 +735,13 @@ task {
quick {
min_query_count: 10
min_duration: 10
max_duration: 40
# 40s could never satisfy min_query_count on a 1B model: measured on an
# iPhone 16 Pro, llm-1b runs about 10.1s per query on the Metal delegate
# (24.3 tok/s), so ten queries need ~101s and the run ended invalid --
# isMinQueryMet false -- on both the CPU and the GPU delegate. A run
# stops as soon as min_duration AND min_query_count are both met, so
# this ceiling only costs time on devices that were failing anyway.
max_duration: 150
}
rapid {
min_query_count: 6
Expand Down Expand Up @@ -777,7 +801,13 @@ task {
quick {
min_query_count: 10
min_duration: 10
max_duration: 40
# 40s could never satisfy min_query_count on a 1B model: measured on an
# iPhone 16 Pro, llm-1b runs about 10.1s per query on the Metal delegate
# (24.3 tok/s), so ten queries need ~101s and the run ended invalid --
# isMinQueryMet false -- on both the CPU and the GPU delegate. A run
# stops as soon as min_duration AND min_query_count are both met, so
# this ceiling only costs time on devices that were failing anyway.
max_duration: 150
}
rapid {
min_query_count: 6
Expand Down Expand Up @@ -837,7 +867,13 @@ task {
quick {
min_query_count: 10
min_duration: 10
max_duration: 40
# 40s could never satisfy min_query_count on a 1B model: measured on an
# iPhone 16 Pro, llm-1b runs about 10.1s per query on the Metal delegate
# (24.3 tok/s), so ten queries need ~101s and the run ended invalid --
# isMinQueryMet false -- on both the CPU and the GPU delegate. A run
# stops as soon as min_duration AND min_query_count are both met, so
# this ceiling only costs time on devices that were failing anyway.
max_duration: 150
}
rapid {
min_query_count: 6
Expand Down
10 changes: 10 additions & 0 deletions flutter/cpp/datasets/ifeval_utils/BUILD
Original file line number Diff line number Diff line change
Expand Up @@ -40,3 +40,13 @@ cc_library(
"@oleander_stemming_library",
],
)

cc_test(
name = "common_test",
srcs = ["common_test.cc"],
linkstatic = 1,
deps = [
":ifeval_utils",
"@com_google_googletest//:gtest",
],
)
11 changes: 8 additions & 3 deletions flutter/cpp/datasets/ifeval_utils/common.h
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,10 @@ inline bool contains_string(const std::string& text,
inline bool ends_with(const std::string& s, const std::string& suf,
unsigned threshold) {
if (s.size() < suf.size()) return false;
std::string a = tolower(s.substr(s.size() - (suf.size() + threshold)));
// The window may be longer than the response itself, in which case it is the
// whole response; subtracting unclamped would wrap and make substr throw.
const std::size_t window = std::min(s.size(), suf.size() + threshold);
std::string a = tolower(s.substr(s.size() - window));
std::string b = tolower(suf);
return threshold == 0 ? a == b : contains_string(a, b);
}
Expand Down Expand Up @@ -142,8 +145,10 @@ inline std::string remove_font_modifiers(const std::string& s) {
}

// skip emphasis/strong/strike/escape chars as long as they're not preceeded
// by an escape character
if ((c == '*' || c == '_' || c == '~' || c == '\\') && s[i - 1] != '\\')
// by an escape character. The first character has nothing before it, so it
// is never escaped -- without the i == 0 guard this reads s[SIZE_MAX].
if ((c == '*' || c == '_' || c == '~' || c == '\\') &&
(i == 0 || s[i - 1] != '\\'))
continue;

// remove heading markers (#) at line starts
Expand Down
Loading
Loading