Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
18 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
93 changes: 93 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,93 @@
# AGENTS.md

This file provides guidance to software agents like Claude Code (claude.ai/code) when working with code in this repository.

## What is Dug

Dug applies semantic web and knowledge graph methods to make research data more findable. It ingests biomedical metadata (e.g. dbGaP study variables), annotates them with ontological terms (via Monarch/SciGraph NLP), normalizes identifiers (via Translator SRI), expands concept graphs (via TranQL), and indexes everything into Elasticsearch for full-text search.

## Setup

```bash
make init # enable git commit-msg hook (run once after clone)
make install # pip install requirements.txt
```

The package sources live under `src/` and PYTHONPATH must include `src/`. The Makefile sets this automatically. If running commands outside make, set:
```bash
export PYTHONPATH=$(pwd)/src
```

Version is single-sourced at `src/dug/_version.py`.

## Testing

```bash
make test # unit tests with coverage
pytest tests/unit # unit tests only
pytest tests/integration # integration tests (requires live services)
pytest tests/unit/test_parsers.py # single test file
pytest tests/unit/test_parsers.py::TestClassName # single test class
```

Integration tests require live Elasticsearch and Redis (see docker-compose). Unit tests run without external services.

## Running the stack

```bash
docker-compose up # starts Elasticsearch, Redis, Neo4J, and Dug search API
```

When connecting to docker services from the host shell, override hosts:
```bash
source .env && export $(cut -d= -f1 .env)
export ELASTIC_API_HOST=localhost
export REDIS_HOST=localhost
```

## CLI usage

```bash
# Crawl (annotate + index) a dataset
dug crawl <file_or_url> -p <parser_type> [-a <annotator>] [-e <element_type>]

# Search
dug search -q "vein" -t concepts
dug search -q "vein" -t variables -k "concept=UBERON:0001638"
```

## Architecture

### Pipeline: crawl → index → search

1. **Parser** (`src/dug/core/parsers/`) — converts raw files (dbGaP XML, CSV, etc.) into `DugElement` and `DugConcept` objects. Parsers are registered via the pluggy hook `define_parsers` in `src/dug/core/parsers/__init__.py`.

2. **Annotator** (`src/dug/core/annotators/`) — takes element descriptions and calls external NLP services (Monarch or SapBERT) to extract `DugIdentifier` (ontology CURIEs). Registered via `define_annotators` hookspec in `src/dug/hookspecs.py`.

3. **Concept Expander / TranQL** (`src/dug/core/concept_expander.py`, `src/dug/core/tranql.py`) — takes identified CURIEs and expands them via TranQL queries to build knowledge graph answers. Results cached in Redis.

4. **Crawler** (`src/dug/core/crawler.py`) — orchestrates parsing → annotation → TranQL expansion for a single file.

5. **Indexer** (`src/dug/core/index.py`) — writes `DugElement`, `DugConcept`, and KG answers into Elasticsearch indices (`concepts_index`, `variables_index`, `kg_index`).

6. **Search API** (`src/dug/server.py`) — FastAPI app exposing `/search`, `/search_var`, `/search_var_grouped`, `/search_kg`, `/search_study`, `/search_program`, and `/program_list` endpoints. Started via uvicorn on port 8181 (default).

### Key types

- `DugElement` (`src/dug/core/parsers/_base.py`) — a single searchable item (e.g. one dbGaP variable).
- `DugConcept` (`src/dug/core/parsers/_base.py`) — an ontology concept that groups elements; holds `DugIdentifier`s and `kg_answers`.
- `DugIdentifier` (`src/dug/core/annotators/_base.py`) — an ontology CURIE with labels, synonyms, and types.

### Plugin system

Dug uses **pluggy** to allow external packages to register new parsers and annotators via setuptools entrypoints (`dug` group). The built-in parsers and annotators are loaded as plugins in `src/dug/core/__init__.py:get_plugin_manager()`.

To add a new parser: implement `Parser = Callable[[Any], Iterable[Indexable]]` and register it in a `define_parsers` hookimpl. Currently supported parsers include `dbgap`, `topmedcsv`, `topmedtag`, `anvil`, `crdc`, `kfdrc`, `sprint`, `bacpac`, `heal-studies`, `heal-research`, `ctn`, `radx`, and others (see `src/dug/core/parsers/__init__.py`).

### Configuration

`src/dug/config.py` defines the `Config` dataclass. All values have defaults and can be overridden via environment variables read in `Config.from_env()`. Key env vars: `ELASTIC_API_HOST`, `ELASTIC_PASSWORD`, `REDIS_HOST`, `REDIS_PASSWORD`. See `.env.template` for the full list.

## Commit convention

This repo uses [Conventional Commits](https://www.conventionalcommits.org/) enforced by the `.githooks/commit-msg` hook. Run `make init` to activate it.
1 change: 1 addition & 0 deletions CLAUDE.md
4 changes: 2 additions & 2 deletions Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
# A container for the core semantic-search capability.
#
######################################################
FROM python:3.13.14-alpine3.24
FROM python:3.13.15-alpine3.24


# Install required packages
Expand Down Expand Up @@ -55,4 +55,4 @@ RUN rm -rf /usr/local/lib/python3*/site-packages/pip \
USER $USER

# Run it
ENTRYPOINT dug
ENTRYPOINT dug
24 changes: 23 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -265,7 +265,7 @@ canonical reference. The current endpoints, grouped by data model type (see [The

| Endpoint | Method | Description |
| ---------------------- | ------ | ---------------------------------------------------------------------------- |
| `/concepts` | POST | Search `DugConcept`s related to a query, with filters/aggregations. |
| `/concepts` | POST | Search `DugConcept`s related to a query, with filters/aggregations/sorting. |
| `/studies` | POST | Search `DugStudy` entities related to a query, or by `parent_ids`/`element_ids`. |
| `/cdes` | POST | Search `DugSection` (CDE set / CRF) entities. |
| `/variables` | POST | Search `DugVariable`/CDE entities, optionally scoped to `parent_ids`. |
Expand All @@ -286,6 +286,28 @@ The `/concepts`, `/studies`, `/cdes`, and `/variables` endpoints (tagged `v2.0`
data-model-backed way to query Dug; the `/search*` and `/dump_concepts`/`/agg_data_types` endpoints predate the Dug
Data Model and are kept for backward compatibility.

### Filtering, aggregating, and sorting

The four `v2.0` endpoints all accept `filters`, `aggs`, and `sort`, each of which names raw Elasticsearch fields.
Field names are passed through to Elasticsearch rather than checked against an allowlist, so the index mappings in
`src/dug/core/index.py` are the source of truth for what's available. A field Elasticsearch rejects comes back as a
`400` naming the reason, not a `500`.

Results are ordered by relevance score unless `sort` is given. `sort` is an ordered list — later keys only break ties
in the earlier ones — and relevance score plus a stable `id.keyword` tiebreaker are always appended, so paging stays
deterministic. Documents missing the sort field go last in both directions.

```shell
curl -X POST http://localhost:5551/studies -H 'content-type: application/json' -d '{
"query": "opioid",
"sort": [{"field": "metadata.Project End Date", "order": "desc"}]
}'
```

The usual pitfall is that `text` fields are not sortable. Sort on a `keyword` subfield (`data_type.keyword`,
`parents.keyword`) instead. Note that `name` and `description` are mapped as `text` with no such subfield in any
index, so they cannot be sorted on today.

## Development

A docker-compose is provided that runs four services:
Expand Down
1 change: 1 addition & 0 deletions requirements_indexinit.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
elasticsearch[async]==8.5.2
2 changes: 1 addition & 1 deletion setup.cfg
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ packages = find:
python_requires = >=3.10
include_package_data = true
install_requires =
dug-data-model
dug-data-model @ git+https://github.com/helxplatform/dug-data-model.git@main
elasticsearch==8.5.2
pluggy
requests
Expand Down
45 changes: 45 additions & 0 deletions src/dug/api_models/request_models.py
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,40 @@ class FilterCriterion(BaseModel):
operator: FilterOperator = Field("eq", description="Comparison operator")
value: Optional[Union[str, int, float, bool, List[Any]]] = Field(default=None, description="The value to filter against")

SortOrder = Literal["asc", "desc"]
SortMode = Literal["min", "max", "sum", "avg", "median"]
MissingPlacement = Literal["_first", "_last"]

# Guardrail on request complexity, in the spirit of Config.aggregate_size_limit
MAX_SORT_FIELDS = 5

class SortCriterion(BaseModel):
field: str = Field(..., description=(
"Elasticsearch field to sort on (e.g. 'metadata.Project End Date'). Must be a "
"doc_values field (keyword/date/numeric/boolean); 'text' fields are not sortable, "
"use their '.keyword' subfield."
))
order: SortOrder = Field("asc", description="Sort direction")
mode: Optional[SortMode] = Field(default=None, description=(
"Which value to sort on for multi-valued fields (parents, programs, variable_list, "
"tags...). Defaults to Elasticsearch behavior: min for asc, max for desc."
))
missing: Optional[MissingPlacement] = Field(default=None, description=(
"Where documents lacking the field are placed. Defaults to '_last' in both directions."
))

@field_validator("field")
@classmethod
def validate_field(cls, v):
v = (v or "").strip()
if not v:
raise ValueError("sort field must not be empty")
# _score and _doc are the only elasticsearch pseudo-fields that can be sorted
# on; letting other underscore names through just yields a confusing ES error.
if v.startswith("_") and v not in ("_score", "_doc"):
raise ValueError(f"'{v}' is not a sortable field")
return v

class SearchElementQuery(BaseModel):
query: str = None
simple_search: bool = False
Expand All @@ -55,6 +89,10 @@ class SearchElementQuery(BaseModel):

aggs: Optional[Dict[str, int]] = Field(default=None, description="Specify fields to aggregate against and the bucket limit")
filters: Optional[List[FilterCriterion]] = Field(default_factory=list)
sort: Optional[List[SortCriterion]] = Field(default_factory=list, description=(
"Ordered list of sort keys, applied before relevance score. Note that fields are "
"ES fields, so subfields like `.keyword` may be required."
))

size: Optional[int] = 100
offset: Optional[int] = 0
Expand All @@ -66,6 +104,13 @@ def drop_empty_strings(cls, v):
return v
return [item for item in v if item not in ("", None)]

@field_validator("sort")
@classmethod
def cap_sort_keys(cls, v):
if v and len(v) > MAX_SORT_FIELDS:
raise ValueError(f"at most {MAX_SORT_FIELDS} sort keys are supported")
return v

class VariableIds(BaseModel):
"""
List of variable IDs
Expand Down
40 changes: 40 additions & 0 deletions src/dug/api_models/response_models.py
Original file line number Diff line number Diff line change
Expand Up @@ -70,3 +70,43 @@ class SectionAPIResponse(DugAPIResponse):
results: List[SectionResponse]


class IndexIngestionMetadata(BaseModel):
index: str = Field(description="Elasticsearch index name")
doc_count: int = Field(description="Total documents in the index")
ingested_at: Optional[str] = Field(
default=None,
description="When the ingestion pipeline last wrote to this index. Null if the "
"index has not been stamped by a pipeline run."
)
index_created_at: Optional[str] = Field(
default=None, description="When the index itself was created"
)


class VariablesIngestionMetadata(IndexIngestionMetadata):
variable_count: int = Field(description="Documents with is_cde false")
cde_count: int = Field(description="Documents with is_cde true")


class IngestionMetadataIndices(BaseModel):
concepts: IndexIngestionMetadata
sections: IndexIngestionMetadata
studies: IndexIngestionMetadata
variables: VariablesIngestionMetadata


class IngestionMappingCounts(BaseModel):
cdes_with_study_mappings: int = Field(
description="CDE sets/CRFs in the sections index carrying a non-empty "
"metadata.study_mappings"
)
variables_with_cde_mappings: int = Field(
description="Variables carrying a non-empty metadata.cde_mapping"
)


class IngestionMetadataResponse(BaseModel):
indices: IngestionMetadataIndices
mappings: IngestionMappingCounts


Loading
Loading