Skip to content

Refine AI test prompts to prefer high-value, non-duplicate coverage - #427

Merged
tninja merged 7 commits into
mainfrom
copilot/fix-ai-generated-tests-again
Jul 9, 2026
Merged

tninja merged 7 commits into
mainfrom
copilot/fix-ai-generated-tests-again

Conversation

Copilot AI commented Jul 9, 2026 •

Copy link
Copy Markdown
Contributor

AI test-generation prompts were biased toward producing too many tests, including low-value overlap. This updates the shared test guidance so generated coverage stays focused on the smallest useful set of distinct behaviors.

  • Shared prompt rules

    • Added a reusable instruction in ai-code-agile.el that tells AI to:
      • prefer a small set of high-value tests
      • cover distinct behaviors, important edge cases, and regressions
      • avoid low-value or duplicate tests
    • Reused the same guidance in ai-code-harness.el so send-time auto-test flows stay aligned with TDD/test-writing flows.
  • Bundled prompt assets

    • Updated packaged prompt files under prompt/ to carry the same high-value/low-duplication guidance:
      • test-after-change.v1.md
      • test-after-change-diagnostics.v1.md
      • tdd.v1.md
      • tdd-diagnostics.v1.md
      • tdd-with-refactoring.v1.md
      • tdd-with-refactoring-diagnostics.v1.md
    • Broke the test-after-change prompt files into shorter lines for readability.
  • Prompt snippets

    • Updated reusable snippets in snippets/ai-code-prompt-mode/ so ad hoc test-generation requests no longer push quantity-first behavior.
    • Replaced the old “generate two tests” shape with guidance that prioritizes the smallest useful set of tests.
  • Generated/autoloaded defaults

    • Updated the checked-in ai-code-autoloads.el default text so the user-visible default suffix stays consistent with the source prompt logic.
  • Regression coverage

    • Added focused ERT assertions to ensure:
      • assembled prompts include the new high-value/non-duplicate guidance
      • bundled prompt assets and snippets stay in sync with that guidance

Example of the new instruction shape:

Prefer a small set of high-value tests. Cover only distinct behaviors,
important edge cases, or regressions that materially increase confidence.
Do not add low-value or duplicate tests.

Copilot AI changed the title [WIP] Fix AI prompt to generate fewer high-value tests Refine AI test prompts to prefer high-value, non-duplicate coverage Jul 9, 2026
Copilot AI requested a review from tninja July 9, 2026 01:27
Copilot finished work on behalf of tninja July 9, 2026 01:27
@tninja
tninja marked this pull request as ready for review July 9, 2026 03:33
@tninja
tninja requested a review from Copilot July 9, 2026 03:37

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e13bf0b2ba

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread ai-code-harness.el
"If any program code changes, run unit-tests and follow up on the test-result (fix code if there is an error). If the tests use random values (for example random numbers or UUIDs), make them reproducible by fixing the random seed or replacing them with deterministic fixtures."
(concat
"If any program code changes, run unit-tests and follow up on the test-result (fix code if there is an error). "
ai-code--high-value-tests-instruction

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Inline this autoloaded default text

Because this defcustom is marked with ;;;###autoload, MELPA/melpazoid-style package builds regenerate an autoload defvar from its initializer and load that autoload file before ai-code-agile has been required. Referencing the private ai-code--high-value-tests-instruction here therefore makes regenerated autoloads depend on an unbound symbol, so package activation can fail with void-variable even though the checked-in autoload was manually expanded; keep the default self-contained or autoload the constant first.

Useful? React with 👍 / 👎.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR refines the repository’s shared AI test-writing guidance to bias generated tests toward a small set of distinct, high-value behaviors (and away from low-value/duplicative coverage). It threads the same instruction through interactive TDD flows, the auto-test harness suffix, and the packaged prompt/snippet assets, with ERT assertions to prevent drift.

Changes:

  • Introduces a reusable “high-value, non-duplicate tests” instruction constant in ai-code-agile.el and reuses it in the harness/TDD suffix assembly.
  • Updates bundled prompt assets (prompt/*.v1.md) and yasnippet prompt snippets to include the same guidance (with readability line breaks where applicable).
  • Adds regression tests to ensure assembled prompts, bundled prompts, and snippets all contain the new guidance.

Reviewed changes

Copilot reviewed 14 out of 14 changed files in this pull request and generated no comments.

Show a summary per file
File Description
ai-code-agile.el Adds shared high-value test guidance constant and injects it into TDD prompt tail assembly.
ai-code-harness.el Reuses the shared guidance in the test-after-change suffix so auto-test flows match TDD behavior.
ai-code-autoloads.el Updates the checked-in autoload default suffix text to match the new guidance.
prompt/test-after-change.v1.md Adds the high-value/non-duplicate guidance and splits into shorter lines for readability.
prompt/test-after-change-diagnostics.v1.md Adds the high-value/non-duplicate guidance while preserving diagnostics instructions.
prompt/tdd.v1.md Extends packaged TDD prompt with the high-value/non-duplicate test guidance.
prompt/tdd-diagnostics.v1.md Extends packaged TDD+diagnostics prompt with the high-value/non-duplicate test guidance.
prompt/tdd-with-refactoring.v1.md Extends packaged TDD+refactoring prompt with the high-value/non-duplicate test guidance.
prompt/tdd-with-refactoring-diagnostics.v1.md Extends packaged TDD+refactoring+diagnostics prompt with the high-value/non-duplicate test guidance.
snippets/ai-code-prompt-mode/unit-tests Updates unit-test snippet to prefer a small, distinct set of tests and avoid duplicates.
snippets/ai-code-prompt-mode/create-tests Replaces “generate two tests” phrasing with high-value/non-duplicate guidance.
test/test_ai-code-agile.el Asserts agile prompt assembly includes the new guidance.
test/test_ai-code-harness.el Asserts harness-generated suffixes include the new guidance.
test/test_ai-code-package-hygiene.el Adds regression checks to ensure autoload defaults and packaged prompt/snippet assets contain the new guidance.

@tninja
tninja merged commit 1745025 into main Jul 9, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Fix: AI generated too many tests. The prompt should be updated to let AI generate less and high value tests.

3 participants