Repository navigation
Conversation
tasks.pbtxt ships six llm tasks; the frontend tables only knew two. - BenchmarkId.allIds omitted llm-3b, llm-3b-instruct, llm-8b and llm-8b-instruct. BenchmarkStore sorts with allIds.indexOf, which returns -1 for a missing id, so those four floated above every other benchmark. - getLocalizedInfo had a case for llm-1b only and its default branch throws. showBenchInfoBottomSheet is the sole caller, and the benchmark set passes benchmarks[0] — one of the -1 sorted tasks — so the LLM set's info button threw 'unhandled task id'. Five of the six llm tasks were affected. - The icon tables mapped the same two ids, so the rest fell back to the generic MLCommons logo. Adds a test that reads assets/tasks.pbtxt and asserts every task in it has an allIds entry, resolvable localized info and its own icon, so a task added to the config fails here rather than under the user's finger. BenchmarkId.llm and .llmInstruct are renamed to .llm1b and .llm1bInstruct now that there are six; all three references were in this change.
A set collapses a cross product of options into fewer controls (#1095), but the card only ever showed the tick count — "1/3 options selected" — and never the benchmarks that count produces. One LLM parameter tick queues two benchmarks through the hidden dataset option set, and nothing on screen said so. That blind spot is where the stored-settings bug in #1180 lived. Config card: - The subtitle states the outcome, "Runs 2 of 6 benchmarks", and the card lists those benchmarks with the backend and delegate each will use. - A set whose options are all off says so instead of reading "0/3". - Options are chips carrying Option.name; the previous rows printed the raw option id, so the LLM set offered "1b", "3b", "8b". - Options are always visible. They are the primary control for a set, and the old layout hid them behind a second disclosure on the same row as the gear. - An option set with max_selected: 1 renders as a single choice and says "pick one". The proto has always had the bound; the UI never expressed it. - Backend and delegate stay per benchmark, as #1164 made them, under one disclosure now labelled Backends. An "Apply to all" control writes one backend to every benchmark in the set that offers it and reads "Mixed" when they differ. A benchmark with a single backend shows plain text and no picker; the delegate picker is omitted when there is no real choice. - The set icon and info sheet come from the set, via BenchmarkSetInfo, rather than from benchmarks[0]. Results card: the header states how many of the set's benchmarks ran, and the ones that did not are marked "Not run" rather than given a result of "N/A". The two cards are extracted into BenchmarkSetCard and BenchmarkSetResultTile, free of BenchmarkState and driven by callbacks, which is what lets them be rendered in a widget test. BackendChoice and the delegate dropdown move to backend_choice.dart alongside a DelegateChoice widget. unit_test/ui/benchmark_set_card_test.dart covers the states and, with SCREENSHOT_DIR set, writes a PNG of each so the design can be reviewed without building for a device.
Adds a screenshot test that captures the config screen at 390x844 with the app's own theme, so the cards can be judged in context — density, how they sit under the GO section, where the fold lands — rather than as isolated widgets. BenchmarkState has a private constructor and late-initialised native dependencies, so BenchmarkStartScreen itself cannot be pumped. The list is the real BenchmarkConfigList holding the real cards; only the chrome around it is reproduced, with the values copied from benchmark_start_screen.dart. To make that possible the loose-benchmark row becomes BenchmarkLooseCard, matching BenchmarkSetCard: free of BenchmarkState, driven by callbacks. The list itself is BenchmarkConfigList, and BenchmarkConfigSection is now just the adapter that wires it to state. The fixture gains a set-less benchmark so the screen shows both shapes.
setSurfaceSize resizes the render surface but leaves MediaQuery reporting the test default of 800x600. Anything measured from MediaQuery therefore laid out for the wrong screen: the GO circle is MediaQuery.width * 0.32, so it came out 256pt across and 272pt tall inside a 390pt-wide phone, and the screenshots made the start screen look far more cramped than it is. Setting tester.view.physicalSize and devicePixelRatio sets both. On a real 390x844 phone the GO section is 140.8pt, 17% of the height, and the list gets 599.2pt — enough for both set cards and the top of the next one.
flutter_test sets debugDisableShadows, which paints every elevated widget as a solid black outline. The GO circle came out with a thick black ring users never see, which reads as a design change next to a screenshot of the current app. The capture helpers turn shadows on, repaint the captured subtree, take the image and restore the flag inside the test body, where the binding's invariant check expects it.
|
MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅ |
anhappdev
marked this pull request as draft
September 26, 2026 13:31
Contributor
|
@mohitmundhragithub and @AhmedTElthakeb please check this UI changes. |
The results screen shows every benchmark with its own icon, but the set card on the start screen listed the same benchmarks behind a plain dot, both in what will run and in the backends panel. The icon replaces the dot in both places; a benchmark the selection leaves out is dimmed in the backends panel, the way the results tile dims one that did not run.
Users found the option chips less clear than the old checkboxes. Pills side by
side ("Offline | Online") read as a pick-one segmented control, and an
unselected chip showed no box, so nothing said it could be ticked.
Each option is now a checkbox with its label, or a radio button when the
option set is bounded to one choice. They keep the one-row layout and the 44pt
tap height, and the rule next to the heading reads "select one or more"
instead of "any".
Collaborator
Author
|
@freedomtan @farook-edev The screenshots in the PR body are updated with your feedback:
|
|
anhappdev
marked this pull request as ready for review
October 7, 2026 12:45
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Stacked on #1180; the diff here is only the commits on top of it.
Start screen: set card
max_selected: 1. A set with nothing selected shows a warning.Results screen: set tile
Fix
Code
BenchmarkConfigSectionwires them toBenchmarkState.task_coverage_testchecks that every task intasks.pbtxthas an id, info and icon. There are also widget tests for the cards, andSCREENSHOT_DIR=<dir>writes phone-size PNGs.Notes
Screenshots: all 13 shipped benchmarks with the LLM set expanded, 390pt @3x. "Before" is #1180's head.