Repository navigation
Add iOS support to mobile_back_litert (CPU + Metal), upgrade to LiteRT 2.2.0 - #1175
Conversation
Adds the Apple half of mobile_back_litert so iPhone runs the benchmarks through the same LiteRT v2 CompiledModel code as Android. The motivation is coverage: no LLM benchmark runs on iOS today (neither the CoreML backend nor the TFLite Apple settings claim llm-*), and iOS gains a vendor-neutral engine next to CoreML for vision/NLP. mlperf_backend_matches_hardware previously returned false off Android; it now serves litert_settings_apple there. The new Apple settings claim the same 12 benchmarks as Android with two delegate choices each, CPU (XNNPACK) and Metal, reusing the Android model URLs verbatim so no new resource can 404 -- listResources walks every delegate choice's model file, so one bad URL would break resource loading for every run. delegate_selected is CPU everywhere for this first landing: Metal is offered but unverified on real hardware, and promoting it is a one-line change per benchmark once a device run confirms it. Metal comes from the prebuilt libLiteRtMetalAccelerator.dylib, which LiteRT dlopens by joining kLiteRtEnvOptionTagRuntimeLibraryDir with the library name. Android does not need that option because jniLibs is on the dlopen search path; iOS has no directory search at all, and the app passes an empty native_lib_path there, so apple_support.h recovers the directory from this binary's own location via dladdr. The environment setup is shared by both pipelines in litert_env.h rather than duplicated. The iOS xcframework needs its own exported symbols list and a 14.0 minimum -- LiteRT's LITERT_MIN_IOS_VERSION, above the 13.1 the sibling backends declare. The orphaned mobile_back_litert/cpp/backend_dummy is replaced by a liblitertbackend stub in mobile_back_tflite's dummy package, where the tflite and coreml stubs already live. Verified: both the Apple and Android configurations of litert_c.cc, single_model_pipeline.cc and llm_pipeline.cc type-check against the pinned LiteRT 2.1.5 headers; the settings parse under protoc; buildifier and clang-format are clean.
Wires the new xcframework into the iOS build the same way the tflite and coreml backends are wired: litert_backend.mk gains ios target/zip variables that fall back to the dummy stub when WITH_LITERT=0 (Xcode fails if a referenced framework is missing), ios.mk builds and unzips it alongside the others, and the Xcode project links and embeds it. The Metal accelerator is downloaded from a pinned 2.1.5 URL rather than through @litert_prebuilts, whose workspace rule points at a floating binaries/latest/ archive. It is embedded but deliberately not linked -- it is dlopened at runtime, and it ships only an ios_arm64 device slice, so putting it on the link line would break simulator builds. The download skips when the file is already present and retries on transient failures, since it now sits on every iOS build's critical path.
Adds a litert row to both the macOS build matrix and the BrowserStack device matrix, alongside tflite and apple, and exports WITH_LITERT from ci_post_clone.sh next to its siblings so the Xcode Cloud default is stated rather than implicit. expected_accuracy gains cpu|LiteRT keys. checkAccuracy resolves '<accelerator>|<backend>' before falling back to the bare accelerator key, and the Apple settings select the CPU delegate, which runs the quantized exports -- the bare 'cpu' intervals were measured on fp32 models, so four of the six vision benchmarks would have failed against them. The intervals are wide until real device numbers exist, matching how the Android LiteRT entries were seeded. The device job timeout goes to 90 minutes: unlike the other iOS rows, the litert row runs llm-1b and llm-1b-instruct, and the go button validates every delegate choice's model file, which means two ~1.2 GB LLM exports per benchmark. That is the same headroom the Android unified job already gets.
Running a vision benchmark on the CPU delegate aborted the app: ERROR: Custom allocation is too small for tensor idx: 320 ERROR: [litert_compiled_model.cc:162] Failed to allocate tensors F external.h:163] Error while inferencing model ResizeInputTensorNonStrict does not allocate. It clears the cached CPU buffer requirements and calls MarkSignatureNeedsAllocation, leaving AllocateTensors to Run. CreateOutputBuffers therefore sees the output tensors at their pre-resize sizes -- for a CPU tensor the buffer requirement is exactly tensor->bytes -- and allocates buffers that are too small. Run then registers those buffers as TFLite custom allocations and calls AllocateTensors, the real shapes finally land, and VerifyCustomAllocationForTensor rejects the undersized allocation. Since a failed Run is LOG(FATAL) in external.h, the app dies. Asking for the output layouts with update_allocation forces the allocation before the buffers are created, so they get their true sizes. The GPU path never hit this because those tensors are not CPU tensors: their requirements come from the accelerator's buffer context rather than tensor->bytes, and they are not host custom allocations. This is not iOS-specific. The vision models export a dynamic batch dimension, so the resize always runs; Android only escaped it because every vision benchmark there defaults to the GPU delegate, leaving the CPU path unexercised.
Device measurements on an iPad mini: Metal ran image_classification_v2 at 77.3 QPS against 44.9 on CPU, and object_detection, image_segmentation_v2, natural_language_processing and image_classification_offline_v2 all completed with valid results. That is the device confirmation the CPU defaults were waiting for, so the six vision/NLP benchmarks now select Metal. super_resolution follows the same pattern but was not in that run; a failed GPU compile falls back to CPU on its own, so the risk is a slower run rather than a broken one. The llm-* benchmarks stay on CPU. They are unmeasured on Apple, and on Android the LiteRT GPU LLM path measured roughly half the speed of LiteRT CPU, so Metal is offered there without being assumed better.
|
MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅ |
The 60->90 job timeout raised earlier was inert. The nick-fields/retry step carries its own timeout_minutes: 60, and that is what fired: both iOS device jobs died at 65m30s (60m timeout + the 300s retry_wait, which elapses even though retry_on_exit_code only covers a trigger rejection). BrowserStack reports a build as running while it is still queued for a device, so the step budget has to cover queue time too. Observed in one run: the apple row spent about 35 minutes queued before a 9.8 minute session. 80 leaves margin under the 90 minute job cap once retry_wait is added. The Android unified job pairs 90 with 85.
The limit observed is 3376 MB on an 8 GB device, so it is not confined to 4 GB-class hardware as previously written, and the CPU delegate hits it too - defaulting llm-* to CPU does not avoid it.
Resolve the conflicts from the stable diffusion PR (#1174): - BUILD: union of both header sets; the SD sources now compile into the iOS framework too, which type-checks clean under __APPLE__. - litert_c.cc: keep the multi-platform comment and fold in the SD pipeline, noting SD is claimed on Android only. - README: the backend is Android arm64 and iOS arm64, claiming SD on Android only; keep the SD file descriptions over the now-stale 'not compiled yet' note. - ios workflow: keep both fixes. max_attempts 8 retries the BROWSERSTACK_ALL_PARALLELS_IN_USE slot rejection, and timeout_minutes 80 is the timeout that actually fires before the job cap.
Bring the iOS settings up from six benchmarks to nine. stable_diffusion now runs on iOS on the CPU delegate, the same exports Android uses. The three models total 1.04 GiB, well inside the memory limit. The pipeline was already platform-neutral; it now builds its environment through CreateLiteRtEnvironment() like the others, so a future GPU choice does not silently miss the Apple accelerator directory. The 1B LLM benchmarks needed a memory fix. The prefill logits buffer is [1, bucket, vocab] fp32, so each bucket doubles it, and picking the smallest bucket at or above the prompt length sent a 2589-token prompt to the 4096 bucket -- a 1.96 GiB buffer that pushed the peak past the per-process limit (measured at 3376 MB on an 8 GB device) and got the app killed with EXC_RESOURCE. On Apple the bucket is now capped at the KV cache length: a larger bucket can never be used, because a prompt longer than the KV cache is rejected outright. That halves the buffer to 0.98 GiB and brings llm-1b to about 2.55 GiB. Overflow tokens decode one at a time, which is slower for long prompts but correct. Android keeps the uncapped choice so its validated throughput does not move. llm-3b and llm-8b are no longer claimed on iOS. Their weights plus the two KV sets are 4.36 GiB and 9.05 GiB before a single logit, so no prefill tuning rescues them; claiming them only offers a benchmark that kills the app. CLAIM_SHARED leaves them to the TFLite fallback. Verified: both settings files parse against the BackendSetting proto, every source type-checks under __APPLE__ and __ANDROID__, and the bucket selection is covered by a test that splices the function out of the source at test time (9 cases, including that Android is unchanged).
The Format & lint job checks clang-format; the wrapped call to GetSuitablePrefillSignature did not match the repo style.
The first iPad mini run got all the way through: the pipeline compiled, Metal registered, the prompt encoded, and all 20 diffusion steps completed at 4.4-5.6 s each. It was then killed at "Image decoding started" with EXC_RESOURCE (high watermark, limit 3376 MB). The VAE decode is the memory peak of the pipeline, and it runs with the text encoder and the 822 MiB diffusion model still compiled, because all three models are built up front and kept for the whole benchmark. The decode needs neither. On Apple they are now released before decoding and rebuilt at the start of the next query, so the peak is the decoder plus its activations instead of the sum of all three. Compiling all three models takes about 1.3 s against a query that takes ~100 s on the CPU, so the cost is negligible. The hooks are null on other platforms, so Android keeps every model resident and its validated throughput does not move. Both sources type-check under __APPLE__ and __ANDROID__.
The download ran `curl -s` with no -f and no retries, then reported success on `[ -f "$log_file" ]` -- true for a zero-byte file. A failed or too-early fetch therefore printed "Device logs downloaded successfully" and uploaded an empty artifact, which reads as "the run produced no output" rather than "the download failed". A litert-iPhone failure had to be diagnosed by inference because of this. Now: -f so an HTTP error is a failure, retries because the log is not always served the instant the session ends, and the success test is -s rather than -f. On failure it writes the reason, session and URL into the file instead of leaving it empty, so the artifact explains itself. Kept non-fatal on purpose: `if-no-files-found: error` on the upload step would otherwise turn a log hiccup into a red device job.
Its only caller is inside the __APPLE__ block, so an unguarded definition is an unused function in an anonymous namespace on every other platform. Both configurations now type-check clean under -Wall.
The CI device run still died during stable_diffusion, ~93 s after the benchmark started, which is about where the diffusion loop ends and the decode begins -- so releasing the text encoder and the diffusion model first was not enough. I have now inferred where the memory goes twice; measure it instead. `os_proc_available_memory()` returns exactly the budget EXC_RESOURCE enforces. LITERT_LOG_MEM logs it at the phase boundaries that allocate: after the three models compile, at query start, after diffusion, after the release, and after the decode. The macro is a no-op off Apple, and the call is a counter read. Note for whoever reads a BrowserStack device log next: it captures only Dart `flutter:` output, no native C++ lines at all, so the absence of backend logs there is not evidence. These numbers have to come from a local `flutter run`.
Releasing the text encoder and the diffusion model before the decode was not enough on its own: the CI device run still died there. free() does not necessarily shrink the process footprint -- libmalloc keeps the pages on its free list, ready to reuse -- and that footprint is exactly what EXC_RESOURCE measures, so a correct-looking release can leave the budget unchanged. This matters here because LiteRT defers AllocateTensors to the first Run, so the decoder's arena, the largest single allocation in the pipeline, is created during the decode rather than at build time. It lands on top of whatever the allocator is still holding from the 822 MiB diffusion model. malloc_zone_pressure_relief(nullptr, 0) reclaims across every zone. Called after the release, and once at backend_create: in a CI sweep this pipeline starts after five other benchmarks have each built and torn down a backend, and it then compiles about 1 GiB of models. Apple-only; both configurations type-check clean under -Wall.
…self The [mem] lines went only to absl LOG(INFO), i.e. native stderr, which never reaches a BrowserStack device-log artifact -- that capture carries os_log entries (Dart's print arrives that way) and no native stderr at all. So a CI memory failure gave pass/fail and nothing to explain it, which is how the last stable_diffusion failure had to be pinned down by timing arithmetic instead of evidence. Logging to both keeps local `flutter run` output unchanged and puts the same numbers in the device log the next red run uploads.
The pressure-relief change worked: the decode now completes. The instrumentation shows it, in MiB still available before the iOS limit on an iPhone 16 Pro: before compiling models 2849 (~527 already used by the app) all three models compiled 1770 (the three models cost 1079) diffusion done 373 (the denoising loop needs ~1405) transient models released 2688 (freeing them recovered 2315) decode done 1279 (the decoder arena needs ~1409) query start (2nd) 541 <- died here So the crash moved rather than disappearing. Releasing was one-sided: the second query rebuilt the encoder and the diffusion model on top of the decoder's ~1.4 GiB arena and had 541 MiB left for a phase that needs ~1405. Each phase needs most of the budget on its own and they never fit together, so the two hooks are replaced by one `set_phase`, which keeps exactly the models the phase needs and releases the rest. Symmetric by construction, which the previous pair was not. Android leaves set_phase null, keeps everything resident, and does not move. Both configurations type-check clean under -Wall.
llm-1b is killed by EXC_RESOURCE on an iPad mini during prefill, even run alone on a freshly launched app, so the prefill cap is not enough on its own. Every number behind that conclusion is arithmetic over model files and tensor shapes, and that arithmetic was already wrong once: the KV cache is four resident sets (prefill in/out plus decode in/out, which TransferKV swaps rather than shares), not the two I had assumed. This measures instead: - LITERT_LOG_MEM at backend_create, model compiled, decode buffers built, prefill buffers built, either side of the prefill Run, and after backend_delete. - The bucket the cap actually selects, with the prompt length and the cap, which nothing reported before -- it is the one input that decides the prefill logits buffer. - The buckets the model exports, kv_cache_max_size, num_kv_layers and vocab_size. - LogBufferSet: the real packed size of each buffer set and the tensor that dominates it, which is what settles the four-KV-set question. LITERT_LOG_NOTE mirrors these to os_log on Apple for the same reason LogAvailableMemory does: a BrowserStack device-log artifact drops native stderr, so absl-only lines are invisible in CI. Also reclaims at stable diffusion's backend_delete, so a benchmark that follows it does not start on top of pages the allocator merely freed -- and logs the budget either side of that reclaim, which shows whether it is worth anything. No behaviour change beyond that reclaim. Both configurations type-check clean under -Wall.
Measured on an iPad mini: selecting the Metal delegate for llm-1b is killed by EXC_RESOURCE (limit 3376 MB) while the delegate is still initialising -- delegate_kernel.cc "Initializing Metal-based API from graph" -- before the model finishes compiling and long before any prefill buffer exists. It consumed the whole ~3.0 GiB budget in about 0.3 s, crashing in __bzero. The graph it materialises carries every prefill signature, and the delegate pays for all of them up front. XNNPACK allocates lazily per signature, which is why the CPU path reaches prefill instead. Nothing in the prefill bucket cap reaches a failure this early, so the choice can never be made to work from here; offering it only hands the user a delegate that kills the app -- the same reason llm-3b and llm-8b are not claimed. CPU was already the default for both, and Android's GPU LLM path ran at about half the speed of LiteRT CPU, so nothing is lost. Android keeps its GPU choice unchanged. Both settings files still parse under `protoc --encode=mlperf.mobile.BackendSetting`.
Measured on an iPad mini (limit 3376 MB), MiB still available: backend_create start 2968 (app baseline 408) model compiled 896 (the weights cost 2072 resident) decode buffers built 894 prefill buffers built 653 prefill inputs written 461 prefill Run -> EXC_RESOURCE, the arena needs more llm-1b does not fit on that device at all. The weights alone take 2072 of 3376 MB, and there is nothing smaller to fall back to: this export publishes exactly one prefill bucket, 1024, and the run above used it on a 381-token prompt. It does run on an iPhone 16 Pro at 9.31 tok/s. So the llm-* settings move to litert_settings_apple_llm.pbtxt and mlperf_backend_matches_hardware appends them only when HasHeadroomForLlm() passes. Claiming them everywhere kills the app on smaller devices; dropping them everywhere gives up the only LLM benchmark any iOS backend offers. The threshold is the weak part: it is bounded below by one measured failure (2968 MiB) and above by one measured success, and 3.25 GiB sits between them. Both branches log the real number and the decision, so any device that runs this narrows it. Two corrections this measurement forces, both to claims made earlier in this PR: - Weights cost 2072 MiB resident against a 1229 MiB file. XNNPACK repacks the q8 weights, so file size is not a proxy for footprint. - There is no large prefill logits buffer. Prefill outputs are 192 MiB over 32 tensors, all KV; the first token comes from the decode signature. The prefill-bucket cap is therefore a no-op for this export, which publishes only one bucket. It is kept because it is correct for an export that publishes several, but it is not what makes anything fit. Both settings files parse standalone and composed under `protoc --encode=mlperf.mobile.BackendSetting`; both configurations type-check clean under -Wall.
The LiteRT backend has to declare 14.0 because litert/cc/BUILD sets LITERT_MIN_IOS_VERSION = "14.0". Leaving the siblings at 13.1 meant the app embedded frameworks with two different minimums for no reason, so raise them all to match. Covers the six apple_xcframework targets the app embeds: the backend bridge, the CoreML backend, the TFLite backend, and the three dummy stubs. Also drops the now-wrong comment on the dummy about not needing the real backend's 14.0. Safe on its own terms: the app's IPHONEOS_DEPLOYMENT_TARGET is already 15.0, so both 13.1 and 14.0 were below the binding constraint, and nothing in ios.mk passes a conflicting --ios_minimum_os. Left alone deliberately: the --macos_minimum_os=13.1 flags in cmdline.mk and mobile_back_apple/dev-utils/Makefile. Those build plain macOS .so dylibs for the desktop CLI, not frameworks the app embeds, and the macOS cmdline target does not link LiteRT at all.
The memory gate is removed. os_proc_available_memory() is not a device property, so it was the wrong thing to branch on: - mlperf_backend_matches_hardware is called repeatedly, and on one iPhone 16 Pro run the value read 3319, 2915, 3045, 3021, 2800, 2643 and 2600 MiB depending on what had already run. The benchmark list therefore depended on when the question was asked. - The threshold needed to sit between one measured failure and one measured success, which put it 9 MiB above an iPhone 16 Pro. That run dropped llm-1b and llm-1b-instruct from a device where they work at 9.31 tok/s. A tested-device allowlist replaces it later; nothing is gated for now, so llm-1b is claimed everywhere again, including the iPad mini where it is known to be killed. The measurements stay: the settings file and README keep the iPad mini trace, the 2072 MiB resident cost against a 1229 MiB model file, the single 1024 prefill bucket, and a note on why the gate was tried and dropped, so the allowlist starts from the evidence. Settings merge back into litert_settings_apple.pbtxt (9 benchmarks, parses under protoc --encode). Both configurations type-check clean under -Wall.
The rewrite added in the previous commit only takes effect once the models carrying it are the ones downloaded. Published alongside the old files rather than over them, under _v2 names, because the base branch and the other open PRs still reference the originals by checksum. sd_diffusion_model_litert_v2.tflite 7bd563de431b8253ea482fdbdd513a47 sd_text_encoder_litert_v2.tflite be6310001157f50bed34750db2c88486 The decoder does not move: it is fp16, has no per-tensor FULLY_CONNECTED weights, and convert.py widened nothing in it, so the published file is byte-identical and keeps its checksum.
Two set_phase conditions changed length when kEncodeAndDiffuse became kEncode and kDiffuse, so one now fits on a single line and the other no longer does. Removing the compile-only probe scaffold also left a double blank line behind.
The litert device job on iPhone 16 Pro was 27 minutes green before the 2.2.0 pin move and stable_diffusion on Metal; it now runs past 80 and is killed mid-run. The kill lands on the step that downloads the BrowserStack device logs, so the run produces no evidence of where the time goes and the slowdown cannot be attributed. Diagnostic, not a fix: 120 keeps the documented margin under the 135 minute job cap minus retry_wait_seconds. Expect this to come back down once the run time is understood.
Renaming the published Metal exports to _v2 moved the download URLs but left
the *_filename custom settings pointing at the old basenames. The app names
the downloaded file after the URL, so the pipeline looked for a file that was
never there:
Could not open '.../liblitertbackend_stable_diffusion_Metal/
sd_text_encoder_litert.tflite'
The model allocation is null/empty
stable_diffusion then failed at its first CompiledModel::Create, on the GPU
attempt and again on the CPU retry, and took the app down with it. It never
reached the memory this branch is about: the device log shows the budget
unchanged at 3012 MiB across both attempts.
decoder_filename is deliberately untouched -- the decoder is fp16, convert.py
widened nothing in it, and it keeps its published name and checksum.
Also reverts the device step timeout to 80 minutes. It was raised to 120 on
the theory that stable_diffusion on Metal had made the suite three times
slower, which the log disproves: five benchmarks pass with valid results in
about a minute each and the run dies the moment stable_diffusion starts. The
80 minute budget was never the problem.
…uilds The device log shows why stable_diffusion was still killed after the models were re-quantised. It was not the size of the models: on an iPhone 16 Pro all three compile to 1368 MiB of the 2885 available. It was releasing them. before compiling models 2885 MiB left all three models compiled 1517 query start, after releasing two of them 1586 <- only 69 recovered diffusion model rebuilt 72 memorystatus: exceeded mem limit: ActiveHard 3376 MB (fatal) ReleaseModel drops the LiteRT objects and ReturnFreeMemoryToOS empties libmalloc's free list, but a GPU-compiled model's weights are Metal buffers libmalloc never owned, so almost nothing comes back. Rebuilding the diffusion model for the next phase therefore stacks a second copy on top of the first, and that is what crosses the limit. So the phases are not installed on the GPU path. The three models are compiled once and stay resident, which they can: 1368 MiB leaves 1517 free for the working set. The CPU path keeps them, where the weights are ordinary allocations the release does return and the three plus a phase do not fit. Not recompiling per query also makes it much faster. One 20-step image on the macOS host goes from 45.5 s to 18.8 s, which is where it was before sharing was turned on -- the throughput sharing costs is bought back by no longer paying for three Metal compiles per query.
The block wrapped in the new if (!use_gpu) guard was reindented by hand. CI pins clang-format 16.0.0 (tools/formatter/Dockerfile); the clang-format in Xcode is 21 and disagrees with it on pre-existing lines, so it cannot be used to check this. Formatted with the pinned version, which now reports the whole repository clean.
App Store Connect rejected build 269:
ITMS-90208: Invalid Bundle - The bundle
Runner.app/Frameworks/LiteRtMetalAccelerator.framework does not support the
minimum OS Version specified in the Info.plist.
The framework's Info.plist is written by hand here, with MinimumOSVersion
14.0. The 2.2.0 accelerator declares 15.0:
$ vtool -show-build libLiteRtMetalAccelerator.dylib
cmd LC_BUILD_VERSION
platform IOS
minos 15.0
so the bundle claimed a release it cannot run on. The 2.1.5 dylib was built
for 14.0, which is why the number was right until the pin moved, and why
nothing caught it before the upload.
Rather than write the new number down, read it back out of the dylib: that is
the only source that cannot drift from the binary actually being shipped. A
dylib without LC_BUILD_VERSION fails the build instead of quietly producing a
plist with an empty version. CFBundleVersion gets the same treatment -- it was
a hardcoded 2.2.0 sitting next to ${backend_litert_version}.
LITERT_MIN_IOS_VERSION is 15.0 as well, not the 14.0 the comments here
claimed -- litert/cc/BUILD says 15.0 in 2.1.5 and 2.2.0 alike -- so the
frameworks bazel builds are raised to match. That is what the app targets
anyway, and every bundle it embeds now declares a minimum it can honour.
Three claims in this README were left behind by the commits that got stable_diffusion onto the delegate: * "the llm-* benchmarks select Metal; stable_diffusion stays on CPU", and "stable_diffusion on Metal does not work at all". It selects Metal, and on the CI iPhone 16 Pro it runs at about 75 s per image against about 98 s on CPU. * The pipeline was described as turning constant tensor sharing off. It turns it on -- that is what the fifth model rewrite is for -- so the paragraph weighing the cost of leaving it off no longer describes anything. * "four blockers", where convert.py fixes five. tools/sd_gpu/README.md has said five since the per-axis rewrite landed. Also names the v5_0 exports where "the models that ship today" had become ambiguous, now that the _v2 models ship alongside them.
|
@freedomtan We now have both LLM and SD running with Metal. You can test it using TestFlight build 270. |
@anhappdev I tested on an iPhone 17 Pro. two issues:
one concern: |
|
To build with Xcode 27.0 on macOS 27.0. One-line change is needed. From fed5f0bd59abbe8d93472cf6b9ce8b2e6c56faaa Mon Sep 17 00:00:00 2001
From: Koan-Sin Tan <koansin.tan@gmail.com>
Date: Mon, 14 Sep 2026 10:53:03 +0800
Subject: [PATCH] set swift version to be 6
With Xcode 27 on macOS 27, to build the coreml backend with
`make WITH_APPLE=1`, we need set the swift version. Either
`copts = ["-swift-version", "5"],`, or `copts = ["-swift-version", "6"],` is fine.
---
mobile_back_apple/cpp/backend_coreml/BUILD | 1 +
1 file changed, 1 insertion(+)
diff --git a/mobile_back_apple/cpp/backend_coreml/BUILD b/mobile_back_apple/cpp/backend_coreml/BUILD
index 271a42d9..57030d3c 100644
--- a/mobile_back_apple/cpp/backend_coreml/BUILD
+++ b/mobile_back_apple/cpp/backend_coreml/BUILD
@@ -56,6 +56,7 @@ swift_library(
srcs = [
"coreml_util.swift",
],
+ copts = ["-swift-version", "6"],
generates_header = True,
deps = [
":apple_frameworks",
--
2.54.0 (Apple Git-157)
|
Thanks for the fix with Xcode 27.0 on macOS 27.0.
I run a test with 17 Pro on BrowserStack and it finished (see updated PR body). What is the error you got?
Yes, I noted that bug too. I can fix it in another PR.
I need to upgrade LiteRT to v2.2.0 to make SD run on Metal. @farook-edev Do you have any concern with upgrading LiteRT to v2.2.0? |
@freedomtan That was a flutter issue, it was caused by the flutter side reading the wrong data and -as a result- canceling the instruct benchmark
@anhappdev I would generally prefer to stick with |
|
@freedomtan to capture the log, it seems to be un-related to our previous one what crashed in Submission run. |
|
@farook-edev let's check we have the same issue to run SD on Android for the GPU accelerator on Android. |
|
for iOS 27, one more patch needed. Previous one only makes it compile. There are runtime errors. It seems, on iOS 27, UIKitCore evaluates runtime issues and raises a debugger trap(brk #0) if the app has not adopted the UIScene lifecycle. When launched in debug mode with LLDB attached, the debugger caught this breakpoint and
diff --git a/flutter/ios/Runner/AppDelegate.swift b/flutter/ios/Runner/AppDelegate.swift
index 62666446..c30b367e 100644
--- a/flutter/ios/Runner/AppDelegate.swift
+++ b/flutter/ios/Runner/AppDelegate.swift
@@ -2,12 +2,15 @@ import Flutter
import UIKit
@main
-@objc class AppDelegate: FlutterAppDelegate {
+@objc class AppDelegate: FlutterAppDelegate, FlutterImplicitEngineDelegate {
override func application(
_ application: UIApplication,
didFinishLaunchingWithOptions launchOptions: [UIApplication.LaunchOptionsKey: Any]?
) -> Bool {
- GeneratedPluginRegistrant.register(with: self)
return super.application(application, didFinishLaunchingWithOptions: launchOptions)
}
+
+ func didInitializeImplicitFlutterEngine(_ engineBridge: FlutterImplicitEngineBridge) {
+ GeneratedPluginRegistrant.register(with: engineBridge.pluginRegistry)
+ }
}
diff --git a/flutter/ios/Runner/Info-Debug.plist b/flutter/ios/Runner/Info-Debug.plist
index 2f766b88..9757f4ed 100644
--- a/flutter/ios/Runner/Info-Debug.plist
+++ b/flutter/ios/Runner/Info-Debug.plist
@@ -43,6 +43,27 @@
<string>LaunchScreen</string>
<key>UIMainStoryboardFile</key>
<string>Main</string>
+ <key>UIApplicationSceneManifest</key>
+ <dict>
+ <key>UIApplicationSupportsMultipleScenes</key>
+ <false/>
+ <key>UISceneConfigurations</key>
+ <dict>
+ <key>UIWindowSceneSessionRoleApplication</key>
+ <array>
+ <dict>
+ <key>UISceneClassName</key>
+ <string>UIWindowScene</string>
+ <key>UISceneDelegateClassName</key>
+ <string>FlutterSceneDelegate</string>
+ <key>UISceneConfigurationName</key>
+ <string>flutter</string>
+ <key>UISceneStoryboardFile</key>
+ <string>Main</string>
+ </dict>
+ </array>
+ </dict>
+ </dict>
<key>UISupportedInterfaceOrientations</key>
<array>
<string>UIInterfaceOrientationPortrait</string>
diff --git a/flutter/ios/Runner/Info-Release.plist b/flutter/ios/Runner/Info-Release.plist
index 30209627..8588fc2c 100644
--- a/flutter/ios/Runner/Info-Release.plist
+++ b/flutter/ios/Runner/Info-Release.plist
@@ -39,6 +39,27 @@
<string>LaunchScreen</string>
<key>UIMainStoryboardFile</key>
<string>Main</string>
+ <key>UIApplicationSceneManifest</key>
+ <dict>
+ <key>UIApplicationSupportsMultipleScenes</key>
+ <false/>
+ <key>UISceneConfigurations</key>
+ <dict>
+ <key>UIWindowSceneSessionRoleApplication</key>
+ <array>
+ <dict>
+ <key>UISceneClassName</key>
+ <string>UIWindowScene</string>
+ <key>UISceneDelegateClassName</key>
+ <string>FlutterSceneDelegate</string>
+ <key>UISceneConfigurationName</key>
+ <string>flutter</string>
+ <key>UISceneStoryboardFile</key>
+ <string>Main</string>
+ </dict>
+ </array>
+ </dict>
+ </dict>
<key>UISupportedInterfaceOrientations</key>
<array>
<string>UIInterfaceOrientationPortrait</string>
|
|
for the error on iOS 27, we can see LLM 1B (Instruct) finished running. It's when it tried to run compute accuracy, which is NOT supposed to do for QUICK mode. after dart_run_benchmark.cc:158 is to LOG information.
|
# Conflicts: # mobile_back_litert/cpp/backend_litert/llm_pipeline.cc
Xcode 27 stopped defaulting the Swift language mode, so `swiftc` refuses to compile `coreml_util.swift` under `make WITH_APPLE=1` without one. The source does not depend on the mode; 6 is what Xcode 27 would choose. Reported by @freedomtan against Xcode 27.0 on macOS 27.0.
On iOS 27 UIKitCore raises a runtime issue when an app has not adopted the scene-based life cycle, which lands as a debugger trap (brk #0) and stops the app under LLDB. Declare a UIApplicationSceneManifest naming FlutterSceneDelegate in both Info plists, and move plugin registration off didFinishLaunchingWithOptions. Under the scene lifecycle the implicit engine is created after that call, so registering there would run against an engine that does not exist yet; FlutterImplicitEngineDelegate hands us the engine once it does. Both protocols are present in the Flutter pinned by CI (3.44.4). Reported by @freedomtan against iOS 27.
remove_font_modifiers tested s[i - 1] to decide whether a modifier was escaped, without excluding i == 0. A response that opens with '*', '_', '~' or '\' therefore indexed the string at SIZE_MAX. Under -O2 the optimiser folds that into a literal dereference of the index, which is the EXC_BAD_ACCESS at 0xffffffffffffffff reported on iOS 27: same function, same +176 offset, and it reproduces on a host clang++ -O2 build. ASan calls it a one-byte buffer underflow at common.h:146. The first character has nothing before it, so it can never be escaped. ends_with had the same shape of bug: it took the last suf.size() + threshold characters after only checking s.size() >= suf.size(). Its one caller passes threshold 3, so any response within three characters of the suffix length wrapped the subtraction and made substr throw std::out_of_range. Clamp the window to the response. Neither is iOS-specific or new -- both date from #1040 and fire on any platform whose codegen happens not to absorb the bad read. The trigger is the content of the generated response, not the run mode: accuracy is computed after every run, for every dataset that has it. Adds common_test with a regression case per bug; it crashes under ASan without the fix. Reported by @freedomtan against iOS 27.
|
|
For f3d44f2, I got the llama 3 1b accuracy numbers on iOS (w/ LiteRT + GPU accelerator)
Look quite reasonable numbers. |
Xcode 27 and iOS 27 support, plus the CI changes that would have caught the breakage before master. The first two commits are cherry-picked from #1175; the rest come from reading the Xcode Cloud logs for builds 272 and 273. - **Xcode 27 build break** (`eea9fe96`) — swiftc refuses `coreml_util` with `emitting module interface files requires '-language-mode'`; Xcode 27 no longer defaults the language mode. The CoreML `swift_library` now passes `-swift-version 6`. The source is mode-agnostic; 6 is what Xcode 27 would pick. - **iOS 27 launch trap** (`fb2444c5`) — UIKit traps when an app has not adopted the UIScene lifecycle. Both Info plists declare `UIApplicationSceneManifest` naming `FlutterSceneDelegate`, and plugin registration moves to `FlutterImplicitEngineDelegate`, since the implicit engine does not exist during `didFinishLaunchingWithOptions`. Both protocols are in the pinned Flutter 3.44.4. - **Intel/Rosetta CI machine** (`9985b710`) — build 273 landed on macOS 26.5.1 where `brew` resolves to the x86_64 install at `/usr/local` under Rosetta 2. No Intel bottles since September 2026, so `protobuf` built from source and died on `clang: error: unsupported argument 'westmere' to option '-march='`. `ci_post_clone.sh` now prefers `/opt/homebrew` and logs the resolved brew. - **Xcode 27 CI coverage** (`f58e0326`) — Xcode Cloud archives with Xcode 27 and runs only on master, so this class of break is invisible until after merge. The two existing legs move from `macos-15` to `macos-26` (newest Xcode 26.6 rather than 26.3), and a new `apple-xcode27` leg builds on the `xcode-27` image — Xcode 27.0 build 27A266a, the same build Xcode Cloud uses. Notes: - The new leg builds and unit-tests only; the test package and BrowserStack stay on the GA image. `xcode-27` is still a GitHub preview with limited capacity, so the leg is `continue-on-error` — a queueing hiccup must not block the BrowserStack jobs that depend on this one. Drop that when the image reaches GA. - Cache keys, the test package and the artifact are now keyed on a new `label` rather than the backend, since two legs build `apple`. Existing check names are unchanged. - Pinning the Xcode Cloud workflow to the macOS 27 image is being done separately in App Store Connect. - The first two issues were reported by @freedomtan.




Adds the Apple half of
mobile_back_litert: iPhone and iPad run all nine benchmarks through the same LiteRT v2 CompiledModel code as Android, on Metal — includingstable_diffusion, which needs the LiteRT 2.2.0 upgrade this branch also carries.Stacked on #1171 — merge that first. Up to date with
litert, including #1174. Supersedes #1178.What it adds
mlperf_backend_matches_hardwarereturned false off Android; it now serves a newlitert_settings_apple.pbtxt. Nine benchmarks claimed, all selecting Metal.apple_xcframework, minimum iOS 15.0 — LiteRT'sLITERT_MIN_IOS_VERSION, and what the prebuilt accelerator declares. Sibling backends and the Flutter framework move up from 13.1 to match; the app already targeted 15.0.libLiteRtMetalAccelerator.dylibpinned at 2.2.0, wrapped in a framework (ITMS-90426rejects a bare Mach-O underFrameworks/; itsMinimumOSVersionis read back out of the dylib, after a stale 14.0 drewITMS-90208) and dlopened via a directory recovered withdladdr— iOS has no dlopen search path andnative_lib_pathis empty there.Runner.entitlementsgainskernel.increased-memory-limitandkernel.extended-virtual-addressing, in a revertible commit of their own.litertrow in both iOS CI matrices;llm-*quick.max_duration40 -> 150.custom_buffer_teardown.patchregenerated with context (its offsets had madedestroy_funcunreachable, so no hardware buffer was released — host teardown 1840 -> 875 MiB), andCreateOutputBuffersnow sizes buffers after resize, which is whyobject_detectioncrashed andimage_classification_v2did not.Why LiteRT 2.2.0
The blocker is
stable_diffusionon Metal. On 2.1.5 the backend emits shader source that does not compile for the diffusion model's int8 weights —use of undeclared identifier 'q0'— withAllowSrcQuantizedFcConvOpsboth on and off, so it is code generation rather than an option we set. The text encoder compiles there; the diffusion model does not.It cannot be patched. The codegen ships only inside the prebuilt
libLiteRtMetalAccelerator.dylib—ml_drift_delegate/in the source archive is TFLite-facing glue with no shader templates, sopatches/*.patchhas no way to reach it. Nor can the accelerator move on its own: a 2.2.0 dylib will not load into a 2.1.5 runtime, or the reverse.So the whole pin moves, and
org_tensorflowwith it — the GPU delegate needsabsl/status:status_macrosfrom the Abseil that TensorFlow brings in (WORKSPACE:186). That coupling is what makes this a toolchain change for every backend and platform rather than a version bump. Patches go seven -> four, plus five for upstream gaps the old pins hid.2.2.0 renames the bug rather than fixing it, which is the next section.
Stable diffusion on Metal
tools/sd_gpu/convert.pyrewrites the published v5_0 exports so the delegate accepts them (fold run-time shapes to constants, dropBROADCAST_TO, squeeze rank-5 tensors, lower a staleSUBversion).A fifth rewrite makes it fit. Constant tensor sharing decides whether weights are materialised or dequantised in-shader; without it the diffusion model needs 4409 MiB against ~2885 available. Enabling it hit
error: use of undeclared identifier 'scale'— the backend picks a scalar weight-dequant template for per-tensor int8 weights without declaring the arguments it references. Single-op models place the fault exactly: per-tensor int8FULLY_CONNECTEDfails; per-axis FC and everyCONV_2Dform compile.With the codegen out of reach, the model is the only side we control: the converter re-expresses those weights as per-axis, repeating the scale they already carry. No weight byte moves, CPU output bit-identical. Diffusion drops to 1913 MiB. Outputs publish as
*_litert_v2.tflite, leaving the originals for the base branch and #1174.Memory
iOS kills a process past a per-process cap (3376 MB on an 8 GB device) and
EXC_RESOURCEcannot be caught. Apple-only; Android untouched.llm-1bneeds 4497 MiB and diffusion 4409; on, 2055 and 1913.set_phaseis therefore CPU-only; on Metal the three compile once and stay resident (1973 MiB of 3013, peak 2588), which also took a 20-step image from 45.5 s to 18.8 s on a macOS host.llm-3b/llm-8bare never offered:llm-1balone costs ~2072 MiB resident from a 1229 MiB file.Results
Nine benchmarks in one process, no crash or relaunch. The 16 Pro is the CI run at
b2f9f166; the two 17 Pros are a BrowserStack run of the same test package at3a0ec360. QPS, exceptllm-*which is tok/s.isMinQueryMet: falseisEarlyStoppingMet: falseisEarlyStoppingMet: falseEvery benchmark selects Metal on all three, with no CPU fallback and no memory kill.
stable_diffusionputs 2732 of 2742 diffusion ops on the GPU: ~75 s/image on the 16 Pro and ~36 s on the 17 Pro, against ~98 s on CPU. Metalllm-1bmeasured 23.5–26.1 tok/s on the 16 Pro, against 9.27 on CPU and 3.46 on the Pixel 10 Pro GPU, TinyMMLU 41% vs 42%. The same BrowserStack build passed on the iPhone 17, 17e and Air too. iPad mini: Metal wins all six valid vision/NLP benchmarks, by 1.6x to 3.0x.Known limits
Interpreter— it does not exercise the fp16 Metal path, and is not a proof for all inputs.stable_diffusionandllm-*are notisResultValid(isMinQueryMet/isEarlyStoppingMetfalse;llm-*would need ~64 queries, ~11 min each). SD already carries an exemption inperformance_result_validity.dart; extending it tollm-*is a scoring-policy call left for a reviewer.llm-1b*on the iPad mini is unmeasured on Metal, and claimed on every iOS device regardless — treat a red iPad run as expected. A tested-device allowlist is the intended fix.iPhone17,1whileexpected_throughput.dartis keyed oniPhone16,2. Adding it would make inert assertions live fortfliteandappletoo — a separate change.mobile_back_apple's job.