Repository navigation
Conversation
…CL binary). Includes updating LiteRT to v2.1.5 and Bazel to v7.7.0, along with compile-time error fixes for Android.
Also fixed LLM benchmark names in binary and potential mask overflow issue
# Conflicts: # .bazelversion # WORKSPACE # flutter/lib/benchmark/benchmark.dart # flutter/lib/ui/home/benchmark_config_section.dart # flutter/lib/ui/settings/resources_screen.dart # flutter/windows/flutter/generated_plugin_registrant.cc # flutter/windows/flutter/generated_plugins.cmake # mobile_back_qti/cpp/backend_qti/SnpeExecutor.cc # mobile_back_qti/cpp/backend_qti/allocator.h # mobile_back_qti/cpp/backend_qti/mlperf_helper.h # mobile_back_qti/cpp/backend_qti/qti_c.cc # mobile_back_qti/cpp/backend_qti/soc_utility.cc # mobile_back_qti/cpp/backend_qti/tflite_c.cc # mobile_back_tflite/cpp/backend_tflite/neuron/neuron_backend.cc # mobile_back_tflite/cpp/backend_tflite/neuron/neuron_builder.h
Combines the latest master (e073a03, per-benchmark backend selection) with the LiteRT/compiled-model work from compiled_model (d93f3d2). Conflict resolutions: - .bazelversion: 7.7.0 (required by LiteRT/TF deps) over master's 6.6.0 - WORKSPACE: LiteRT 2.1.5 archive + separate org_tensorflow source repo. Master's xcode26_compat.patch was dropped from the litert archive: it targets tensorflow/lite/kernels/elementwise.cc, a path LiteRT 2.1.5 does not have (it uses tflite/). Needs re-porting for Xcode 26 builds. - QTI/neuron sources: kept master's refactors (SnpeExecutor rename, backend_utils, Genie support) and re-applied compiled_model's tensorflow/core/platform/logging.h -> absl/log/log.h swap, extending it to files master added that compiled_model never saw. - mobile_back_qti/.../tflite_c.cc: accepted master's deletion. - Dart + generated Windows plugin files: kept master's versions (newer dart formatter style, current plugin set). Reverted two dev-only changes that commit 530df21 carried alongside the dependency update, as they are regressions in a combined branch: - flutter/assets/tasks.pbtxt: restored the 6 llm-{1b,3b,8b}[-instruct] task definitions it deleted. - flutter/lib/backend/loadgen_info.dart: restored the real tokenThroughput computation (was hardcoded to double.infinity).
|
MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅ |
The C++ and bazel format jobs failed on code carried over from compiled_model, which never ran CI. Master is clean on both. clang-format: compiled_model's tensorflow/core/platform/logging.h -> absl/log/log.h swap left the include blocks unsorted in 6 files; one loop in llm_pipeline.cc also needed rejoining. buildifier (6.0.1, matching CI): reformatted 5 files and fixed the three lint warnings that fail the job: - WORKSPACE loaded and called python_init_rules twice, verbatim - mobile_back_tflite BUILD loaded litert_gpu_accelerator_prebuilts but never used it (also unused on compiled_model, so nothing is lost) - tensorflow_source_rules.bzl print() is deliberate (it surfaces the patch script's output), so it is marked buildifier: disable=print rather than removed Verified with the exact commands CI runs: clang-format reports 0 errors repo-wide and buildifier exits 0.
All four build jobs failed fetching @eigen_archive: Cannot find patch file: external/xla/third_party/eigen3/eigen_ios.patch tf-eigen.patch has two hunks. The second patches third_party/xla/third_party/eigen3/workspace.bzl to add patch_file = ["//third_party/eigen3:eigen_ios.patch"]. The first creates that patch file, but its +++ line said third_party/eigen3/eigen_ios.patch while its own diff --git header said third_party/xla/third_party/eigen3/eigen_ios.patch. The +++ line wins, so the file landed at the TF root instead of inside the XLA tree. This was harmless on master: TF 2.18.0 still has a root third_party/eigen3, and that root copy is the one its build evaluates, so the label resolved inside @org_tensorflow and found the file. TF 6d40c20c, which LiteRT 2.1.5 pins, removed the root third_party/eigen3 entirely -- eigen now lives only under third_party/xla. So XLA's copy is what gets evaluated, and its label resolves inside @xla, where the file was never created. Point the +++ line at the path the hunk's own header already names.
Android builds failed compiling SdExecutor.cc: QnnApiHelpers.hpp:65:43: error: unknown type name 'float32_t' float32_t is a NEON type. The vendor header uses it without including arm_neon.h, relying on it arriving transitively; TensorFlow's headers used to provide that and LiteRT's do not. The header is downloaded at build time so it cannot be patched in-tree -- include arm_neon.h from SdExecutor.h instead, guarded on NEON being available.
backend_utils.cc failed to compile: error: implicit instantiation of undefined template 'std::basic_istringstream<char>' Only <iosfwd> was reaching it. TensorFlow's headers used to pull in <sstream> transitively and LiteRT's do not, so include it directly. qti_c.cc uses std::stringstream the same way and is in the same target, so it gets the include too. The other files the scan flagged (coco_gen.cc, ifeval.cc, mmlu_gen.cc) compiled fine in this build, so they still get it transitively and are left alone.
iOS builds failed before compiling anything: plisttool, line 39: %interpreter_args% SyntaxError: invalid syntax rules_apple's plisttool is a py_binary whose bootstrap stub is generated from a template. TensorFlow pins rules_apple 3.5.1, which predates the rules_python that XLA now brings in, so the newer template's %interpreter_args% placeholder is never substituted and the stub is not valid Python. Declare the same versions LiteRT 2.1.5 pins for this dependency set, before the TensorFlow workspace macros so they win: rules_apple 3.22.0, rules_swift 2.9.0, apple_support 1.23.1, bazel_features 1.43.0. apple_support is deliberately 1.23.1 rather than TensorFlow's 1.24.5 -- per LiteRT, that is the highest version without missing LC_UUID, DEVELOPER_DIR or SDKROOT on macOS Tahoe, which matters here since CI runs Xcode 26.3. All four sha256 values verified against the archives.
Android: the neuron backend compiles tflite_c.cc and llm_pipeline.cc, which now use the LiteRT compiled-model API, but its deps were never updated to match //mobile_back_tflite/cpp/backend_tflite:tflite_c, so the build failed with undeclared inclusions for ~28 litert headers. Mirror that target's litert deps. iOS: the rules_apple bump got past plisttool, and the build now fails generating @coremltools//:mlmodel_proto: DataStructures.proto: File not found. The .proto files import each other by bare filename while protoc runs with -Iexternal/coremltools. TensorFlow declares coremltools without a fix for this; LiteRT declares the same archive with a patch_cmds that rewrites the imports to be repo-relative. Declare it ahead of TensorFlow's with that patch. Both projects' coremltools.BUILD are byte-identical, so the vendored copy is TensorFlow's file unchanged.
…size Android: the neuron target's deps are fixed but its copts were still missing what //mobile_back_tflite/cpp/backend_tflite:tflite_c sets, so single_model_pipeline.cc failed: no matching function for call to 'TfLiteInterpreterOptionsAddDelegate' no known conversion from 'TfLiteDelegate *' to 'TfLiteOpaqueDelegate *' The shared sources pass a TfLiteDelegate, so this target needs the same -UTFLITE_USE_OPAQUE_DELEGATE and -fpermissive. iOS: coremltools protos now generate and the build gets as far as compiling LiteRT itself, where the CoreML delegate fails: no member named 'resize' in 'google::protobuf::RepeatedField<float>'; did you mean 'Resize'? Patch that call site. The float16 branch below it operates on a std::string, where lowercase resize is correct, so it is left alone; util.cc's resize is on a std::vector and is also fine. Patch verified to apply against the v2.1.5 source.
The neuron target failed with an undeclared inclusion: 'external/org_tensorflow/tensorflow/lite/c/common.h' Our own sources are migrated to tflite/c/common.h; this include comes from @neuron_delegate's neuron_delegate.h. MediaTek's delegate is built against TensorFlow's TF Lite C API rather than LiteRT's, so the header still has to come from @org_tensorflow. Put the dep on the target that exposes that header so it propagates to consumers. Note this leaves the neuron backend pulling TensorFlow's tflite/c:common alongside LiteRT's equivalent. Getting MediaTek's delegate onto LiteRT headers would be the real fix if those collide at link time.
The neuron target now compiles but fails to link: ld.lld: error: duplicate symbol: TfLiteIntArrayCreate ... ~20 more TfLite* C API symbols Depending on @org_tensorflow//tensorflow/lite/c:common to satisfy neuron_delegate.h pulled in a second implementation of the TF Lite C API alongside LiteRT's, so every shared symbol was defined twice. Rewrite that one include to tflite/c/common.h via patch_cmds on the neuron_delegate archive and depend on @litert//tflite/c:common instead, so there is a single TF Lite C API in the .so. The two headers declare the same API; only APUWareUtilsApi.h and neuron_delegate.h are consumed from that archive, and only the latter has the include.
Patching neuron_delegate.h in the archive fixed our link but broke the archive's own build: neuron_delegate.h:28:10: fatal error: 'tflite/c/common.h' file not found (from @neuron_delegate//neuron/java/src/main/native:native) That archive is built against TensorFlow throughout -- its targets depend on tensorflow/lite/c, core/api, delegates/utils, kernels and java/jni -- so it cannot be moved onto LiteRT by rewriting one include, and its JNI .so is a separate shared library that is fine staying as it is. The only thing our code needs from that header is TfLiteDelegate, and LiteRT declares it identically. So generate a copy of just that header with the include repointed, and expose that instead. Quoted includes search the includer's directory before -I paths, so it shadows the vendor header for this package only, the archive builds unmodified, and the backend .so still has a single TF Lite C API.
The generated-header approach did not shadow anything: generated files
live under bazel-out, so the includer's-directory rule never finds them
and -Iexternal/neuron_delegate still won. The vendor header was used
unchanged and the build was back to:
missing dependency declarations for
'external/neuron_delegate/neuron/neuron_delegate.h'
'external/org_tensorflow/tensorflow/lite/c/common.h'
Nothing else in this target provides the tensorflow/lite/c/common.h
path, so instead of fighting include order, add a forwarding header at
that path which includes LiteRT's tflite/c/common.h, and put its
directory on the include path via includes = ["tf_compat"]. The vendor
header resolves it unambiguously, the archive stays unmodified, and the
backend still links a single TF Lite C API.
The forwarding header could not win: Bazel passes -iquote external/org_tensorflow for this target, and -iquote is searched before anything includes = [...] can add, so the vendor header's "tensorflow/lite/c/common.h" always resolved to TensorFlow's copy. Putting the forwarding header where -iquote . would find it would mean creating tensorflow/ at the repo root, which would shadow that path for TensorFlow's own sources too. Every route through that include is blocked: depending on @org_tensorflow//tensorflow/lite/c:common duplicates the TF Lite C API at link, supplying the header chain as files is blocked because tensorflow/lite/core/c exports only to //tensorflow/lite:__subpackages__, and patching the archive breaks its own TensorFlow-based targets. So stop routing through it. Only single_model_pipeline.cc used the vendor header, under MTK_TFLITE_NEURON_BACKEND, and all it needs is TfLiteDelegate, which LiteRT declares identically. Declare the delegate's API against LiteRT's header instead. The declarations are byte-identical to the archive's, verified by diff; the header records that they must stay in sync if the pinned sha256 is bumped.
mobile_back_tflite, mobile_back_pixel, mobile_back_qti, mobile_back_apple and flutter/cpp return to their master state and build against @org_tensorflow//tensorflow/lite again. The LiteRT compiled-model LLM pipeline moves to a new mobile_back_litert backend that claims only the llm-* benchmarks, only on Android, with fallback_policy FALLBACK_FILL_GAPS so every other benchmark stays on the TFLite fallback. liblitertbackend is listed before libtflitebackend in list.in, so LLM routes to LiteRT when both are present.
The builtin py_binary leaves %interpreter_args% unfilled in its launcher under this workspace's rules_python, so every pbtxt2header action fails with a SyntaxError. Loading py_binary from @rules_python fixes it, the same way the branch did before the mobile_back_litert restructure.
litert_c.cc includes litert_settings_android.h, the checked-in wrapper over the pbtxt2header output, following the tflite backend's pattern; it was missing from the new backend. Both iOS jobs fail compiling org_tensorflow's CoreML delegate: the protobuf bundled with TF 6d40c20c renames RepeatedField::resize() to Resize(). Apply the same one-line patch already used for LiteRT's copy of the delegate.
flutter/cpp:utils pulls tensorflow/lite/tools/evaluation:utils and with it TensorFlow's TF Lite C API implementation, which duplicates the one @litert provides, failing the liblitertbackend.so link. llm_pipeline.cc's include of flutter/cpp/utils.h was dead - nothing from it is referenced.
…r inputs
The three macOS bazel jobs (Android 35m08s, iOS tflite 14m53s, iOS apple
14m48s) were on the same actions/cache plus --disk_cache setup the Linux build
just moved off, and they get nothing from it. Two runs of the Android-macOS job
over the same commit range:
run 98382919294 Cache not found 2715 processes: 281 internal, 2434 local
run 98417618611 Cache restored (166 MB) 2715 processes: 281 internal, 2434 local
Identical action counts and zero hits from a populated cache, so this is action
key instability rather than a storage problem. It cannot be the Linux cause:
these workflows have no setup-gcloud step, so nothing puts a per-run uuid on
PATH, and both runs already had --incompatible_strict_action_env. The runner
image was also identical (macos-15-arm64/20260727.0256), which rules out an
image rotation.
Switching to GCS alone would therefore reproduce exactly what #1171 hit -- a
correctly configured cache that hits nothing -- so this does both halves.
Storage: both workflows drop actions/cache and point bazel at GCS, Android
under bazel-cache/macos-android and both iOS backends under bazel-cache/macos-ios
so each backend warms the cache for the other. Since these jobs, unlike the
Linux one, do not otherwise need a GCS credential, BAZEL_CACHE_ARG is chosen at
runtime and falls back to a disk cache when the secret is absent, so
pull_request runs from forks keep working.
Evidence: a new bazel-cache-fingerprint.sh records an allowlisted environment
and toolchain fingerprint before the build and a sha256 manifest of every
fetched external repo after it, stores both in GCS keyed by run number, and
diffs against the previous run of the same job. That is the method that found
the Linux cause, where exactly one environment value and two of ~240k files
differed. It is portable to macOS, which has no sha256sum, no nproc and no GNU
find -printf, and it can never fail a build.
Note that adding setup-gcloud puts a per-run random path on PATH here too. That
is only harmless because .bazelrc now sets --incompatible_strict_action_env.
Ports the dormant stable-diffusion pipeline to the LiteRT 2.1.5 CompiledModel API and wires it in, closing the last benchmark gap on this backend. Targets `litert`, not master. ### How it works - Three models (text encoder, diffusion, decoder) share one `litert::Environment`; 43 invocations per query at the shipped `num_steps: 20`. Buffers are created once and reused across the denoising loop. - **Inputs are bound by signature name, never by position.** The signature keys are ordered alphabetically, which does not match tensor order: the text encoder reads `(positions, tokens)` and the diffusion model reads `(context, latent, timestep_embedding)`. Binding positionally would swap latent with context - both float32, so it would produce a plausible but wrong image rather than an error. A missing name is a hard failure that logs the names actually found. - CPU only. The shipped exports are `dynamic_int8` (encoder, diffusion) and `dynamic_fp16` (decoder), aimed at CPU/XNNPACK; a GPU choice would need dedicated fp32 exports. Despite the filenames, nothing is quantized at the graph boundary - encoder inputs are int32, everything else float32. The delegate is labelled `CPU`/`cpu` accordingly, rather than copying the TFLite entry's `NNAPI`/`npu` labels, which are cosmetic (that code attaches no delegate). - No new model exports and no `tasks.pbtxt` change: the existing v5_0 TFLite SD models are plain flatbuffers and load as-is. All four checksums were verified against the downloaded files. ### Test coverage, and its limit Stable diffusion joins the LLM benchmarks in `quickRun`. In accuracy mode loadgen ignores `min_query_count` and runs the whole set - here 150 sequential 20-step generations plus 150 CLIP scoring passes and a 1.71 GB ground-truth download, which is hours on a CPU against a 90-minute job. `quickRun` keeps the time-bounded performance phase. The consequence is worth stating plainly: **CI exercises the pipeline end to end but cannot catch a wrong-but-plausible image**, because nothing scores the output. Name-based binding plus hard failure on a missing name is what guards correctness instead. The QTI-only gate in `canRunBenchmark` is widened for LiteRT by inverting the condition rather than adding `|| litert`, so QTI behaviour is bit-identical and SD does not start running on the Samsung and MediaTek jobs, which would add ~2.8 GB of downloads and an uncapped CPU run to three more device budgets. ### Defects fixed in the dormant sources Carried over from the TFLite twin, all pre-existing: `exit(-1)` on inference failure now returns `MLPERF_FAILURE`; the per-query invoker leak is gone; the unbounded scan for a token terminator is bounded at 77 and the token tail zero-padded, so no stale data from a previous sample reaches the model; the cumulative-alpha index is bounds-checked; the embeddings loader validates its reads before allocating. Dead code dropped: `get_tensor_index_by_name`, two unused `get_timestep_embedding` implementations, and a `main()` behind `__TEST_BPE__`. Note these files were byte-identical to `mobile_back_tflite`'s copies, so they now diverge - the same defects remain in the TFLite backend and are worth a follow-up there. ### Verification Confirmed on a Pixel 10 Pro (run 33053477420). Stable diffusion ran end to end on `liblitertbackend`: the pipeline initialised, all six tensors resolved by signature name with no binding failure, and both samples completed the full 20-step denoising loop - 111.31 s mean latency per image (111.71 s p90), 0.00895 QPS, 2 samples in 229.9 s. `isResultValid` is false only because the min-query rule (2 vs 128) is unmet under the time-bounded run, exactly as the LLM benchmarks behave. Two incidental fixes rode along, both surfaced by this work: - `checkAccuracy` dereferenced `accuracyRun` with a null check operator, which is null in `quickRun`. The LLM benchmarks never hit it only because their expected-accuracy maps are empty and return early. The run mode's own `doAccuracyRun` flag is now threaded through from the call site, so the check is skipped only when the mode genuinely has no accuracy phase - rather than whenever the field happens to be null, which would silently stop checking benchmarks that should have one. - The BrowserStack trigger retry went from 2 attempts to 8. The device jobs have no `concurrency:` group, so two boards at once exhaust the account's parallel budget and every loser fails outright with `BROWSERSTACK_ALL_PARALLELS_IN_USE` - no session created, nothing run. That cost four device-job failures on this branch that looked like code failures. `retry_on_exit_code: 9` is specifically the trigger failure, so this only lengthens how long a job waits for a slot; everything else still fails fast. A concurrency group would address the cause rather than the symptom, but it changes scheduling for every board and belongs in its own change. Build side: the sources type-check against the pinned LiteRT headers and compile under Code Analysis with `WITH_LITERT=1`; `dart analyze`, `dart format`, `clang-format` and `buildifier` are clean; the settings block parses under `protoc`; the four model checksums match the md5s of the actually-downloaded files.
|
@farook-edev to add some changes including Gemma 4 ones. |
|
@farook-edev text encoder, unet/diffusion model (50 iterations), VAE decoder. Let's try to enable the UNET one on GPU accelerator first. It seems no re-conversion is needed. Let's try it. |
…d time (#1173) Makes the GCS bazel cache from #1171 actually hit across CI runs, and rations the BrowserStack device tests to the account's 2-session limit. ## bazel cache Two things stopped the cache from ever hitting: - **Linux** inherited `PATH` into every action key. `setup-gcloud` puts a per-run `$RUNNER_TEMP/<uuid>` on `PATH`, so every key changed every run — the cache held ~19k entries and hit 7 of 4497 actions. - **macOS** had three jobs sharing one `actions/cache` key over one path. Only one could write it, so the Android job kept restoring the iOS job's cache and never saved its own. Fixes: - `--incompatible_strict_action_env` in `.bazelrc` pins the action `PATH`. Repo rules still see the real environment. - Pin loadgen's build-date stamps in `WORKSPACE` (the only non-reproducible fetched file). - Linux uses the GCS remote cache. macOS keeps a per-config disk cache **in front of** GCS (fast local hits, GCS for cold cases). - Every cache key is now per configuration, so the collision can't recur. Fork PRs with no GCS secret fall back to disk-only. - Fix the CocoaPods key (was hashing an untracked `Podfile.lock`, collapsed to a constant). Result — all jobs now hit 100% of cacheable actions (zero local): | job | before | after | |---|---|---| | Linux build | ~67 min | **33 min** | | macOS Android | 35 min | **23 min** | | iOS apple / tflite | ~15 min | ~15 min (durability, not speed) | ## BrowserStack The account allows 2 parallel sessions; runs were failing with `BROWSERSTACK_ALL_PARALLELS_IN_USE`. - Per-PR `concurrency` group on the four device jobs: a new commit cancels the previous run's job for the same device. Never workflow-wide — builds always finish. - `max-parallel: 1` on those matrices, so one run can't exceed the budget alone. - Device-job `timeout-minutes` 60 → 135, so the built-in retry (`max_attempts: 2`) can actually run instead of being killed after the first attempt. --- Once merged, #1171 can drop its `ci: move the Linux bazel cache to GCS` commit and rebase.
# Conflicts: # .bazelrc # .github/workflows/android-build-test-linux.yml # WORKSPACE
Bazel runs genrules through BAZEL_SH with a fixed PATH of C:\msys64\usr\bin;C:\msys64\bin plus the Windows directories, so the Strawberry Perl on the runner image is invisible to them. The TFLite acceleration/configuration schema genrules shell out to perl and the runner installs msys2 with the base package set only, which ships sed but not perl, so the build died with "perl: command not found". Every previous green Windows run restored a warm bazel disk cache and resolved the whole graph from it (0 local actions), so the genrule never ran. This run was the first with all caches expired, which exposed it.
…terializes them The Android Linux job copies build products straight out of bazel-bin. Bazel 7 defaults to --remote_download_outputs=toplevel, so only the outputs of targets named on the command line are downloaded. Two of the copied files are not top-level outputs -- they are merely srcs of a cc_library that is: flutter/android/commonlibs/lib_arm64/libc++_shared.so (qti, samsung) external/neuron_delegate/.../libtensorflowlite_neuron_jni.so (mediatek) so when their actions come back as remote cache hits nothing writes them to disk and the copy fails with "cp: cannot stat". Request the producing targets explicitly. This was latent until adaa361 made the GCS cache actually hit: before that every action ran locally and produced the files as a side effect. The QTI step failed first; mediatek and samsung would have followed. Also switch the mediatek jni path to ${BAZEL_LINKS_PREFIX}bin/ like every other entry, instead of hardcoding bazel-bin/.
|
gemma4 model files available here |
|
|
I found this on quantization here |
|
If are willing to download models and put them to your device manually, then you can try https://github.com/mlcommons/mobile_app_open/actions/runs/34861220129/artifacts/10356801955. Otherwise, we can wait for the automatically downloading one. |
The llm-gemma-e2b and llm-gemma-e2b-instruct benchmarks pointed at local:///gemma4/... with empty checksums, so the models had to be pushed onto the device by hand and nothing verified what landed there. Upload the four files to gs://mlperf-mobile-public/litert/gemma4/ -- the same bucket llm-1b already uses for llama_q8_ekv3072_litert.tflite -- and point every model_file entry at the https URL with its md5. The basenames are kept as-is because the model_filename, embedder_filename, per_layer_embedder_filename and tokenizer_filename custom_settings reference them verbatim; the gemma4/ subdirectory keeps the generic names (model_quantized.tflite and friends) from colliding in the flat litert/ namespace. Uploaded with parallel_composite_upload_enabled=False: composite objects carry no md5 in GCS metadata, which would have made the checksums here impossible to verify server-side and broken the pattern set by the existing llama object.
llmGemmaE2b and llmGemmaE2bInstruct are in BenchmarkId.allIds, so the integration test enumerates and runs them, but neither benchmarkExpectedAccuracy nor benchmarkExpectedThroughput had an entry. checkAccuracy asserts the per-benchmark map is non-null before it reaches the "skip when there is no expected value" branch, so both benchmarks failed with "missing expected accuracy map" on every device job. Accuracy reuses the existing empty _llm map: the LLM benchmarks have no reference accuracy yet, and these jobs run in quickRun mode where accuracy_run is null anyway. Throughput gets its own _llmGemma rather than reusing _llm, because _llm's 2..50 tok/s bounds do not hold here -- llm-gemma-e2b measured 1.7 tok/s on Pixel 10 Pro, and the instruct variant reported 63.6 and 405.6 tok/s for the same model. Registering the backend with an empty per-device map keeps the non-null assertion satisfied while leaving the value comparison skipped, so CI stops failing without pinning bounds to numbers that are not yet trustworthy.
Gemma E2B measured on BrowserStackBoth Gemma benchmarks now download from The only run where all four completed is build 1035. Build 1036 (
These are not performance numbers. Every run reports Two things worth a look. The instruct throughput is not measuring generation rate. Same model and device as the non-instruct row, yet 405 tok/s against 1.67. It also rises as the device gets slower: Pixel 9 Pro 136 s → 63.6, Pixel 10 Pro 421 s → 405.6. 421 s at 405 tok/s implies ~170k tokens from 2 queries. Something is off in the token accounting on the IFEVAL path. The Gemma benchmarks exceed BrowserStack's idle timeout, which is what the two red device jobs are.
Most of that is the 4.7 GB download plus the LiteRT GPU compile ( Raising |
|
@farook-edev is it possible to check pipeline with the Gemma 4 E4B |
CPU benchmark using cmdline on a desktop computer with 16GB+ memory should be doable, I'm not sure it'll work on Android because of ram limitations (E2B model was hitting 10GB peaks and 8GB sustained RAM usage). P.S. the memory usage is for the entire phone, where the system uses approximately 4-6GB of memory. |
@farook-edev Google AI Edge Gallery already allows running Gemma 4 E4B IT on some Android and iOS devices, please check it out. |
|
…T 2.2.0 (#1175) Adds the Apple half of `mobile_back_litert`: iPhone and iPad run all nine benchmarks through the same LiteRT v2 CompiledModel code as Android, on Metal — including `stable_diffusion`, which needs the **LiteRT 2.2.0** upgrade this branch also carries. Stacked on #1171 — merge that first. Up to date with `litert`, including #1174. Supersedes #1178. ## What it adds - `mlperf_backend_matches_hardware` returned false off Android; it now serves a new `litert_settings_apple.pbtxt`. Nine benchmarks claimed, all selecting Metal. - New `apple_xcframework`, minimum iOS 15.0 — LiteRT's `LITERT_MIN_IOS_VERSION`, and what the prebuilt accelerator declares. Sibling backends and the Flutter framework move up from 13.1 to match; the app already targeted 15.0. - Metal comes from a prebuilt `libLiteRtMetalAccelerator.dylib` pinned at 2.2.0, wrapped in a framework (`ITMS-90426` rejects a bare Mach-O under `Frameworks/`; its `MinimumOSVersion` is read back out of the dylib, after a stale 14.0 drew `ITMS-90208`) and dlopened via a directory recovered with `dladdr` — iOS has no dlopen search path and `native_lib_path` is empty there. - `Runner.entitlements` gains `kernel.increased-memory-limit` and `kernel.extended-virtual-addressing`, in a revertible commit of their own. - A `litert` row in both iOS CI matrices; `llm-*` `quick.max_duration` 40 -> 150. - Two fixes that are not iOS-specific: `custom_buffer_teardown.patch` regenerated with context (its offsets had made `destroy_func` unreachable, so no hardware buffer was released — host teardown 1840 -> 875 MiB), and `CreateOutputBuffers` now sizes buffers after resize, which is why `object_detection` crashed and `image_classification_v2` did not. ## Why LiteRT 2.2.0 The blocker is `stable_diffusion` on Metal. On 2.1.5 the backend emits shader source that does not compile for the diffusion model's int8 weights — `use of undeclared identifier 'q0'` — with `AllowSrcQuantizedFcConvOps` both on and off, so it is code generation rather than an option we set. The text encoder compiles there; the diffusion model does not. It cannot be patched. The codegen ships only inside the prebuilt `libLiteRtMetalAccelerator.dylib` — `ml_drift_delegate/` in the source archive is TFLite-facing glue with no shader templates, so `patches/*.patch` has no way to reach it. Nor can the accelerator move on its own: a 2.2.0 dylib will not load into a 2.1.5 runtime, or the reverse. So the whole pin moves, and `org_tensorflow` with it — the GPU delegate needs `absl/status:status_macros` from the Abseil that TensorFlow brings in (`WORKSPACE:186`). That coupling is what makes this a toolchain change for **every** backend and platform rather than a version bump. Patches go seven -> four, plus five for upstream gaps the old pins hid. 2.2.0 renames the bug rather than fixing it, which is the next section. ## Stable diffusion on Metal `tools/sd_gpu/convert.py` rewrites the published v5_0 exports so the delegate accepts them (fold run-time shapes to constants, drop `BROADCAST_TO`, squeeze rank-5 tensors, lower a stale `SUB` version). A fifth rewrite makes it *fit*. Constant tensor sharing decides whether weights are materialised or dequantised in-shader; without it the diffusion model needs **4409 MiB** against ~2885 available. Enabling it hit `error: use of undeclared identifier 'scale'` — the backend picks a scalar weight-dequant template for **per-tensor int8** weights without declaring the arguments it references. Single-op models place the fault exactly: per-tensor int8 `FULLY_CONNECTED` fails; per-axis FC and every `CONV_2D` form compile. With the codegen out of reach, the model is the only side we control: the converter re-expresses those weights as per-axis, repeating the scale they already carry. No weight byte moves, CPU output bit-identical. Diffusion drops to **1913 MiB**. Outputs publish as `*_litert_v2.tflite`, leaving the originals for the base branch and #1174. ## Memory iOS kills a process past a per-process cap (3376 MB on an 8 GB device) and `EXC_RESOURCE` cannot be caught. Apple-only; Android untouched. - Sharing is on. Off, `llm-1b` needs 4497 MiB and diffusion 4409; on, 2055 and 1913. - **Releasing a GPU-compiled model does not return its memory** — the weights are Metal buffers libmalloc never owned. On device, releasing two of three models recovered 69 MiB of ~1268, and rebuilding one stacked a second copy and crossed the limit. `set_phase` is therefore CPU-only; on Metal the three compile once and stay resident (1973 MiB of 3013, peak 2588), which also took a 20-step image from 45.5 s to 18.8 s on a macOS host. - `llm-3b`/`llm-8b` are never offered: `llm-1b` alone costs ~2072 MiB resident from a 1229 MiB file. ## Results Nine benchmarks in one process, no crash or relaunch. The 16 Pro is the CI run at `b2f9f166`; the two 17 Pros are a BrowserStack run of the same test package at `3a0ec360`. QPS, except `llm-*` which is tok/s. | benchmark | 16 Pro | 17 Pro | 17 Pro Max | | | --- | --- | --- | --- | --- | | image_classification_v2 | 93.7 | 105.9 | 102.7 | valid | | object_detection | 247.5 | 285.5 | 268.9 | valid | | image_segmentation_v2 | 145.2 | 194.4 | 186.6 | valid | | natural_language_processing | 50.7 | 61.9 | 60.4 | valid | | super_resolution | 13.5 | 16.0 | 16.2 | valid | | image_classification_offline_v2 | 99.0 | 140.3 | 137.9 | valid | | stable_diffusion | 0.0133 | 0.0280 | 0.0274 | `isMinQueryMet: false` | | llm-1b | 25.0 | 38.1 | 37.1 | `isEarlyStoppingMet: false` | | llm-1b-instruct | 25.0 | 38.5 | 33.5 | `isEarlyStoppingMet: false` | Every benchmark selects Metal on all three, with no CPU fallback and no memory kill. `stable_diffusion` puts 2732 of 2742 diffusion ops on the GPU: ~75 s/image on the 16 Pro and ~36 s on the 17 Pro, against ~98 s on CPU. Metal `llm-1b` measured 23.5–26.1 tok/s on the 16 Pro, against 9.27 on CPU and 3.46 on the Pixel 10 Pro GPU, TinyMMLU 41% vs 42%. The same BrowserStack build passed on the iPhone 17, 17e and Air too. iPad mini: Metal wins all six valid vision/NLP benchmarks, by 1.6x to 3.0x. ## Known limits - The converter's equivalence check is one seeded input set per model through the CPU `Interpreter` — it does not exercise the fp16 Metal path, and is not a proof for all inputs. - `stable_diffusion` and `llm-*` are not `isResultValid` (`isMinQueryMet` / `isEarlyStoppingMet` false; `llm-*` would need ~64 queries, ~11 min each). SD already carries an exemption in `performance_result_validity.dart`; extending it to `llm-*` is a scoring-policy call left for a reviewer. - `llm-1b*` on the iPad mini is unmeasured on Metal, and claimed on every iOS device regardless — treat a red iPad run as expected. A tested-device allowlist is the intended fix. - No iOS device job asserts throughput, for any backend: BrowserStack reports `iPhone17,1` while `expected_throughput.dart` is keyed on `iPhone16,2`. Adding it would make inert assertions live for `tflite` and `apple` too — a separate change. - No CoreML or ANE path: LiteRT v2 does not expose one. CoreML stays `mobile_back_apple`'s job.
# Conflicts: # .github/workflows/ios-build-test-macos.yml
|
|
@farook-edev: qwen <- converted model no fully-delegated to GPU.
ifeval: will be pushed soon. |




Adds
mobile_back_litert: a backend for Android and iOS that runs every benchmark on the LiteRT 2.2.0 CompiledModel API - the LLM benchmarks on a dedicated LLM pipeline, stable diffusion on the stable-diffusion pipeline, the vision/NLP benchmarks on a single-model pipeline. Android runs on the GPU and CPU delegates; iOS, added by #1175 (merged into this branch), runs on Metal.How it works
llm_pipelinedrives alitert::CompiledModelwith explicitTensorBuffers for prefill and decode.llm-1bandllm-1b-instructdefault to the GPU delegate with a dedicated GPU model export (llama_q8_ekv3072_litert.tflite). GPU compilation failure falls back to CPU automatically.single_model_pipelineruns the vision/NLP benchmarks onlitert::CompiledModel. The GPU accelerator is the default (fp32 model exports, automatic CPU fallback); the CPU delegate choice runs the int8 exports on XNNPACK. Inputs and outputs stage through host buffers so the mlperf harness keeps stable pointers.stable_diffusion_pipelineruns stable diffusion (added in Run stable diffusion on the LiteRT backend #1174). CPU only on Android: the shipped exports are dynamic_int8 (text encoder, diffusion) and dynamic_fp16 (decoder), aimed at CPU/XNNPACK, so a GPU choice would need dedicated fp32 exports. iOS runs it on Metal from rewritten exports (see iOS below). The sub-models bind by name - the CompiledModel signature order is alphabetical, not positional.libtflitebackend, so LLM defaults to LiteRT; on Pixel devices the pixel backend still outranks it for the vision benchmarks.Backend coexistence
fallback_policyis replaced by aclaim_policychain: per benchmark, the app walks the backends in priority order and stops after the first claimant that is notCLAIM_SHARED.CLAIM_SHARED, so users can pick any capable backend per benchmark. Defaults still follow the priority order, so CI runs are unchanged.Build and packaging
WITH_LITERTdefaults to 0, and the unified test APK, release APK/AAB, and Android CLI builds request it explicitly.l, and the Pixel letter changes fromgtop(unified suffix is nowqsmplt).android-apk-litert-<n>).WITH_LITERT=1: this backend is first-party code and stays analyzed, unlike the proprietary vendor SDKs.iOS (#1175)
stable_diffusion,llm-1bandllm-1b-instruct- through the same CompiledModel code, all defaulting to Metal with CPU selectable.llm-3b/llm-8bare not offered (llm-1balone is ~2 GB resident against a ~3.4 GB per-process cap), and the Gemma benchmarks are not claimed on iOS.libLiteRtMetalAccelerator.dylib2.2.0, wrapped in a framework and dlopened from the backend's own directory. Minimum iOS moves from 13.1 to 15.0 for every backend and the Flutter framework (the app already targeted 15.0).Runner.entitlementsgainskernel.increased-memory-limitandkernel.extended-virtual-addressing.tools/sd_gpu/convert.py(published as*_litert_v2.tflite; same weights, CPU output bit-identical) to get past Metal codegen faults and fit the memory cap. On Metal the GPU-compiled models stay resident, because releasing them does not return their memory.custom_buffer_teardown.patchis regenerated so hardware buffers are actually released,CreateOutputBufferssizes buffers after resize (theobject_detectioncrash), andllm-*quick.max_durationgoes from 40 to 150 intasks.pbtxt.llm-1b25.0 / 38.1 tok/s,stable_diffusion~75 s / ~36 s per image, all vision/NLP results valid. The full table, the memory analysis and the known limits are in Add iOS support to mobile_back_litert (CPU + Metal), upgrade to LiteRT 2.2.0 #1175.Merged from
compiled_model.so, ported to the CompiledModel API (see above). The stable-diffusion pipeline followed in Run stable diffusion on the LiteRT backend #1174, also ported to CompiledModel. The dummy-backend scaffolding was imported unwired and later removed in Add iOS support to mobile_back_litert (CPU + Metal), upgrade to LiteRT 2.2.0 #1175. The old iOS/Windows settings variants and the dead iOS delegate branches were dropped; iOS came back through Add iOS support to mobile_back_litert (CPU + Metal), upgrade to LiteRT 2.2.0 #1175 with a newlitert_settings_apple.pbtxt.Toolchain changes
@litert2.2.0;org_tensorflowpinned to the commit LiteRT expects. Add iOS support to mobile_back_litert (CPU + Metal), upgrade to LiteRT 2.2.0 #1175 moved it from 2.1.5: the 2.1.5 Metal accelerator could not compile the diffusion model, and the prebuilt accelerator must match the runtime version. The bump applies to every backend and platform.rules_apple/apple_supportpinned to LiteRT's versions; small patches for the pinned TF revision.minSdkVersion). Building at 33 made absl referencebacktrace(), which Android 11 devices lack, so the backends failed to dlopen there.CI
Device jobs always test the unified APK; the integration test decides the backend. On the Pixel 10 Pro it pins every benchmark to the LiteRT backend (
deviceBackendOverride), so that job covers LiteRT vision and LLM (GPU) plus stable diffusion (CPU) end to end.The Pixel 9 Pro job keeps the defaults: vision on the pixel backend,
llm-1b/llm-1b-instructon LiteRT.Measured on the Pixel 10 Pro job at
1b9e045(LiteRT 2.1.5, before Add iOS support to mobile_back_litert (CPU + Metal), upgrade to LiteRT 2.2.0 #1175) - all nine benchmarks onliblitertbackend,All tests passed. Accuracy comes from the CI subset in quick mode, so it is only meaningful against TFLite on the same run, not as a real accuracy score.LLM CPU baselines: 4.9-8.0 tok/s (Pixel 9 Pro) and 7.97-8.14 tok/s (Pixel 10 Pro). Expected throughput intervals are still deliberately wide (smoke test, not a perf gate); they can be tightened now that the GPU and LiteRT-vision numbers have landed.
The unified device jobs got a 90-minute budget and a longer pre-run wait in the test: with the GPU delegate default the LLM benchmarks carry two ~1.2 GB models, doubling the download and checksum-validation time.
Vendor jobs (qti/samsung/mtk) skip LLM until those backends support it.
tf_nnapi_no_mmap_sharing.patch: the rebuilt TFLite passed the model file's fd to NNAPI drivers. SELinux rejects that for models on external storage, so NNAPI fell back to CPU. The patch restores master's behavior; Pixel EdgeTPU throughput matches master again.Windows build moved from Google Cloud Build to GitHub Actions (
windows-build-test.yml).MSVC needed three fixes: hide host Android NDK env from repo rules; LiteRT's
is_msvc = Truerewrite forxla/tsl; protobuf's--define=protobuf_allow_msvc=true.The Linux Android job uses a GCS bazel cache (same pattern as the Xcode Cloud build) instead of actions/cache, whose 10 GB quota was churned by buildkit blobs and whose keys never re-save.
Workflow downloads are pinned, sha256-verified, and https-enforced.
iOS gets a
litertrow in both the build matrix and the BrowserStack matrix (iPhone 16 Pro). After master's matrix rework (deps: support Xcode 27 and iOS 27 #1181) the row builds onmacos-26.iOS builds ride out known flakes: on a remote-cache error the build retries without the cache;
pod installbacks off on CDN 429s.The legacy Cloud Build trigger can be disabled once the new workflow is trusted.
Review notes
llm-*tasks intasks.pbtxtare restored, andtokenThroughputno longer returnsdouble.infinity. Please confirm neither was deliberate.