Skip to content

Latest commit

 

History

History
170 lines (119 loc) · 6.41 KB

File metadata and controls

170 lines (119 loc) · 6.41 KB

Contributing to Jabberjay

Thank you for taking the time to contribute! Jabberjay's primary goal is to be the one-stop shop for synthetic voice detection models, so every addition — whether a new model, a bug fix, or a documentation improvement — directly advances that mission.

Table of Contents


Ways to contribute

  • Add a model — the highest-impact contribution; see Adding a new model below
  • Report a bug — open a bug report
  • Request a model — open a model request
  • Improve documentation — fix typos, clarify examples, extend the README
  • Write tests — increase coverage, especially for edge cases and error paths

Development setup

You will need uv and just installed.

# 1. Fork the repository on GitHub, then clone your fork
git clone https://github.com/<your-username>/Jabberjay.git
cd Jabberjay

# 2. Install all dependencies including dev tools
just install

# 3. Install pre-commit hooks (runs lint + format automatically on every commit)
uv run pre-commit install

# 4. Verify everything works
just check   # lint + format check + type check
just test    # run the test suite

Work on develop — do not target main directly.


Making changes

  1. Create a branch from develop:
    git checkout develop
    git checkout -b feat/my-change
  2. Make your changes.
  3. Run just fix to auto-format and lint.
  4. Run just check — all checks must pass before opening a PR.
  5. Run just test — all tests must pass.
  6. Update CHANGELOG.md under [Unreleased].
  7. Open a pull request against develop.

Adding a new model

New models are the most valuable contribution. The bar for inclusion is:

Requirement Detail
Licence Apache 2.0 or MIT only
Task Binary bonafide / spoof classification
Availability Publicly available weights (HuggingFace Hub preferred)
Input Raw audio waveform (not pre-extracted features)

Step-by-step

1. Add the model value to the Model enum (src/Jabberjay/Utilities/enum_handler.py):

class Model(Enum):
    ...
    MyModel = "MyModel"

2. Create a run.py in a new directory under src/Jabberjay/Models/MyModel/:

If your model fits the standard transformers audio-classification pipeline, delegate to the shared run_pipeline() helper — it loads and caches the pipeline for you (see Utilities/pipeline.py):

# src/Jabberjay/Models/MyModel/run.py
import numpy as np

from Jabberjay.Utilities.pipeline import run_pipeline
from Jabberjay.Utilities.types import PredictionScore

_MODEL_ID = "author/model-id-on-huggingface"
_TARGET_SR = 16_000


def predict(y: np.ndarray, sr: float) -> list[PredictionScore]:
    return run_pipeline(_MODEL_ID, y, sr, "MyModel", sampling_rate=_TARGET_SR)

If your model needs custom loading (a torch.nn.Module, a non-pipeline HuggingFace class, a .joblib/.pth file, etc. — see the Spectra*, RawNet2, or Classical model families for examples), wrap the load in Jabberjay.Utilities.model_cache.cached_loader instead of calling .from_pretrained()/load() directly inside predict(). Every model in Jabberjay caches its loaded weights in-process so repeated detect() calls for the same model don't pay the reload cost — a new model that skips this would be the one exception and would noticeably slow down multi-call use (e.g. examples/run_all.py):

from Jabberjay.Utilities.model_cache import cached_loader

@cached_loader(maxsize=4)
def _load_model(device: str) -> MyModelClass:
    return MyModelClass.from_pretrained(_MODEL_ID).eval().to(device)

If the model uses non-standard labels, normalize_pipeline_scores() calls normalize_label() internally — check src/Jabberjay/Utilities/label_normalizer.py and extend _BONAFIDE_SUBSTR / _SPOOF_SUBSTR (substring matching) or _BONAFIDE_EXACT / _SPOOF_EXACT (exact matching) if needed.

3. Add a handler and match case in src/Jabberjay/jabberjay.py:

# In detect():
case Model.MyModel:
    return self._mymodel_handler(y=y, sr=sr)

# New method:
def _mymodel_handler(self, y: np.ndarray, sr: float) -> DetectionResult:
    import Jabberjay.Models.MyModel.run as MyModel
    scores = MyModel.predict(y=y, sr=sr)
    logger.debug(f"MyModel predictions: {scores}")
    return self._result_from_scores(scores, Model.MyModel)

4. Update tests:

  • add the new enum value to TestEnums.test_model_members in tests/test_enums.py
  • add a predict() unit test in tests/test_models.py (mock all model loading — the suite runs offline)
  • add a handler test to TestDetectHandlers in tests/test_jabberjay.py

5. Update CHANGELOG.md under [Unreleased] with the model name, HuggingFace link, and dataset.

6. Update the docs — add a row to the models table in README.md and a section in docs/models.md (and docs/cli.md if it changes the CLI choices).

7. Verify everything passes:

just fix
just check
just test

Submitting a pull request

  • Target the develop branch, not main
  • Fill in the pull request template fully
  • Link any related issues with Closes #<issue>
  • Ensure just check and just test both pass locally before opening the PR

Code style

  • Formatter: black (line length 88) — run just format
  • Linter: ruff — run just lint
  • Type checker: ty — run just type-check
  • All at once: just fix (auto-fix) or just check (check only)

Logging uses loguru: logger.info() for model loading, logger.debug() for inference details. Do not use print() inside library code — only in main().