Keywords: Conversational Memory · Discord Export Analysis · Semantic Search · Hybrid Retrieval · Timeline Reconstruction · Retrieval Evaluation · FastAPI
▶ Try the live demo — runs entirely in your browser against a bundled sample export. No token, install, or server required. Search it, reconstruct a timeline, evaluate retrieval quality, and compare embedding backends right from the page.
NeuralLog is a local-first research system for transforming Discord conversations into a searchable operational memory layer. It applies to any community that thinks out loud in chat — engineering teams, study groups, research labs, hobby servers, support channels, open-source projects — where decisions, rationale, answers, and context are distributed across long-lived channels rather than captured in formal documentation. Engineering discussion is used throughout as a concrete worked example, but nothing in the system is specific to it.
This repository combines two layers:
neurallog/: the Python retrieval application, including ingestion, indexing, search, evaluation, timeline reconstruction, and a local web interfaceDiscordChatExporter.*: the upstream Discord export tooling retained as the ingestion substrate for producing compatible JSON exports
The present implementation should be understood as an applied retrieval system rather than a production knowledge platform. Its purpose is to support exploration of how chat archives of any kind can be indexed, queried, evaluated, and iteratively improved using lightweight local infrastructure.
Communities generate substantial knowledge in conversational systems, yet much of it remains difficult to recover once it is buried in chat history. NeuralLog investigates whether Discord export archives can be restructured into a usable memory layer through message normalization, chunk-based indexing, configurable embeddings, hybrid retrieval, and query-conditioned timeline reconstruction. The current system supports local experimentation through both a command-line interface and a FastAPI-backed web application, with evaluation utilities for comparing retrieval behavior across embedding configurations.
Across many kinds of communities, consequential decisions are made in informal channels: debugging threads and integration reviews in an engineering team, paper discussions and experiment notes in a research group, troubleshooting exchanges in a support channel, planning and lore in a hobby or gaming server. These artifacts are valuable, but they are rarely organized for retrospective analysis. (The examples in this document are drawn from an engineering channel, simply because that is the bundled sample export.)
NeuralLog is motivated by the following question:
Can chat history be transformed into a practical retrieval layer for reconstructing decisions, diagnosing past problems, and recovering context?
This repository explores that question with a deliberately compact architecture that can be executed locally and extended incrementally.
NeuralLog currently supports:
- ingestion of Discord JSON exports produced in the style of DiscordChatExporter
- normalization of message metadata, authorship, timestamps, attachments, and references
- chunk-based grouping of temporally adjacent discussion windows
- configurable embedding backends:
hashsentence-transformersopenai
- optional persistent embedding caching in SQLite
- vector search using
inmemoryand optionalfaissbackends - lightweight hybrid ranking that combines semantic similarity with lexical term coverage
- one-shot search directly over an export file
- timeline reconstruction from retrieved evidence
- retrieval evaluation against labeled query sets
- side-by-side comparison across embedding specifications
- a local web UI for interactive experimentation
The present prototype does not yet provide:
- multi-source ingestion from Git, ROS logs, experiment notebooks, or deployment systems
- large-scale persistence and serving infrastructure
- GPU-oriented indexing beyond optional FAISS usage
- LLM-native summarization or agentic reasoning workflows
- formal access control, multi-user deployment, or production hardening
neurallog/
api.py
cache.py
chunking.py
cli.py
discord_ingest.py
embeddings.py
evaluation.py
index.py
models.py
services.py
web/
frontend/
src/
dist/
examples/
sample-discord-export.json
sample-evaluation.json
DiscordChatExporter.Cli/
DiscordChatExporter.Core/
DiscordChatExporter.Gui/
The current processing pipeline is:
DiscordChatExporter JSON
->
message normalization
->
temporal chunking
->
embedding generation
->
vector indexing
->
hybrid retrieval
->
timeline reconstruction / evaluation / API responses
The retrieval stage combines semantic similarity with lexical evidence. In practice, this hybrid ranking matters for queries containing domain-specific terms — a jargon acronym, a product name, a person's handle, an error code — where pure semantic similarity may over-generalize toward merely adjacent concepts. (For example, in an engineering channel, terms like EKF, AMCL, or odom drift should not collapse into nearby topics such as mapping or SLAM.)
NeuralLog targets Python 3.11+.
Install the base package:
python3 -m pip install -e .Optional extras:
python3 -m pip install -e ".[faiss]"
python3 -m pip install -e ".[embeddings]"Embedding backends:
hash: dependency-light baseline intended for portability and fast experimentationsentence-transformers: local semantic embeddings using models such asall-MiniLM-L6-v2openai: hosted embeddings using models such astext-embedding-3-small
By default, embedding outputs are cached in neurallog-embeddings.sqlite3. The cache path can be overridden with --embedding-cache-path, or disabled by passing an empty string.
Start the FastAPI server from the repository root:
python3 -m uvicorn neurallog.api:app --reload --host 127.0.0.1 --port 8000Then open:
http://127.0.0.1:8000
The local web application supports:
- Discord export ingestion
- one-shot search over an export
- timeline reconstruction
- retrieval evaluation
- backend comparison
- direct export creation through the bundled DiscordChatExporter CLI when available
The following figures illustrate the current local interface and representative retrieval output.
If you want to run the React frontend in development mode:
cd frontend
npm install
npm run devThe development frontend expects the API at http://127.0.0.1:8000.
PYTHONPATH=. python3 -m neurallog.cli ingest path/to/export.jsonInstalled-script form:
neurallog ingest path/to/export.jsonPYTHONPATH=. python3 -m neurallog.cli search-export \
examples/sample-discord-export.json \
"Why was localization unstable in March?" \
--limit 3With sentence-transformers:
PYTHONPATH=. python3 -m neurallog.cli \
--embedding-backend sentence-transformers \
--embedding-model all-MiniLM-L6-v2 \
--embedding-batch-size 32 \
search-export examples/sample-discord-export.json \
"Why was localization unstable in March?" \
--limit 3With OpenAI embeddings:
OPENAI_API_KEY=your_key_here \
PYTHONPATH=. python3 -m neurallog.cli \
--embedding-backend openai \
--embedding-model text-embedding-3-small \
--embedding-batch-size 64 \
search-export examples/sample-discord-export.json \
"Why was localization unstable in March?" \
--limit 3PYTHONPATH=. python3 -m neurallog.cli timeline-export \
examples/sample-discord-export.json \
"IMU drift and EKF stability" \
--limit 4PYTHONPATH=. python3 -m neurallog.cli --index-path neurallog.index.json search \
"Summarize EKF tuning decisions" \
--limit 5PYTHONPATH=. python3 -m neurallog.cli \
--embedding-cache-path /private/tmp/neurallog-embeddings.sqlite3 \
evaluate examples/sample-discord-export.json examples/sample-evaluation.json \
--limit 3Returned summary metrics include:
- mean precision@k
- mean recall@k
- mean reciprocal rank
Running the command above (--limit 3) against examples/sample-discord-export.json
and the 3-query labeled set in examples/sample-evaluation.json:
| Embedding backend | mean precision@3 | mean recall@3 | mean reciprocal rank |
|---|---|---|---|
hash (dependency-light) |
0.556 | 0.722 | 1.000 |
These figures come from a deliberately tiny fixture (8 messages, 3 queries) and are
meant as a reproducible sanity baseline, not a benchmark. The perfect reciprocal rank
reflects that the first relevant chunk is always retrieved first on this set; precision
and recall are bounded by the small number of labeled messages per query. Run
compare-backends (below) with sentence-transformers or openai specs to reproduce
the same metrics against learned embeddings on your own export.
PYTHONPATH=. python3 -m neurallog.cli \
--embedding-cache-path /private/tmp/neurallog-embeddings.sqlite3 \
compare-backends examples/sample-discord-export.json examples/sample-evaluation.json \
--spec hash \
--spec sentence-transformers:all-MiniLM-L6-v2 \
--skip-unavailable \
--limit 3Each spec may be provided as:
hashsentence-transformers:model_nameopenai:model_name
Given an export containing discussion fragments such as:
2026-03-10 Maya: IMU drift got much worse after the pool test. EKF position estimate diverges after 90 seconds.
2026-03-12 Alex: We changed the MPU-9250 low-pass filter constants and reduced some of the high-frequency noise.
2026-03-14 Maya: Finished tuning EKF covariance values. Localization is still shaky during turns but much better.
2026-03-17 Jordan: After the covariance tuning and filter changes, localization stayed stable for the full test run.
NeuralLog can answer a query such as:
PYTHONPATH=. python3 -m neurallog.cli search-export \
examples/sample-discord-export.json \
"Why was localization unstable in March?" \
--limit 3Representative output:
{
"results": [
{
"chunk_id": "channel-eng-9f6bf828c801",
"score": 0.4331988897471611,
"channel_name": "systems-integration",
"start_time": "2026-03-14T20:41:00+00:00",
"end_time": "2026-03-14T20:41:00+00:00",
"participants": ["Maya"],
"preview": "[Maya] Finished tuning EKF covariance values. Localization is still shaky during turns but much better.",
"message_ids": ["1003"]
},
{
"chunk_id": "channel-eng-70b8dcb93382",
"score": 0.29285113019775794,
"channel_name": "systems-integration",
"start_time": "2026-03-17T21:05:00+00:00",
"end_time": "2026-03-17T21:05:00+00:00",
"participants": ["Jordan"],
"preview": "[Jordan] After the covariance tuning and filter changes, localization stayed stable for the full test run.",
"message_ids": ["1004"]
}
]
}This behavior illustrates the intended retrieval pattern: chunks explicitly containing the target concept are promoted above merely adjacent topics.
Run the API:
python3 -m uvicorn neurallog.api:app --reload --host 127.0.0.1 --port 8000Primary endpoints:
GET /GET /healthPOST /ingest/discord-exportPOST /searchPOST /timelinePOST /workflow/search-exportPOST /workflow/timeline-exportPOST /workflow/export-discordPOST /workflow/evaluatePOST /workflow/compare-backends
Example request:
curl -X POST http://127.0.0.1:8000/workflow/search-export \
-H "Content-Type: application/json" \
-d '{
"export_path": "examples/sample-discord-export.json",
"query": "Why was localization unstable in March?",
"limit": 3,
"config": {
"backend": "auto",
"embedding_backend": "hash",
"embedding_model": "all-MiniLM-L6-v2",
"embedding_batch_size": 32,
"embedding_cache_path": "neurallog-embeddings.sqlite3"
}
}'NeuralLog expects Discord exports in the JSON format produced by DiscordChatExporter. The exporter source is retained in this repository because it remains the operational starting point for the current ingestion workflow.
If Discord history is collected from a personal account, note that automating user accounts may violate Discord's Terms of Service. When possible, use a bot token scoped only to channels that your application is authorized to access.
This repository incorporates the upstream DiscordChatExporter codebase as an ingestion dependency and foundation for Discord archive generation. Relevant references to the upstream project remain in the bundled source and build metadata, including links to:
https://github.com/Tyrrrz/DiscordChatExporter
NeuralLog should therefore be interpreted as a downstream research system built on top of that export substrate, rather than as an unrelated implementation from first principles.
This repository includes an MIT license; see License.txt. The bundled DiscordChatExporter-derived components also reference MIT licensing in the retained upstream materials.
NeuralLog is currently best suited for:
- local experimentation
- architecture validation
- retrieval quality studies
- conversational-memory demonstrations
- iterative tuning of chunking, embeddings, and ranking behavior
The hash backend remains useful as a portable baseline, but it should be treated as a convenience model rather than the target retrieval ceiling. For higher-quality search behavior, sentence-transformers is the recommended default local backend.
Planned directions include:
- stronger hybrid retrieval and reranking strategies
- more rigorous benchmark datasets across domains
- multi-source ingestion beyond Discord (other chat exports, issue trackers, notes, and logs)
- narrative summarization and root-cause synthesis
- knowledge graph construction across systems, teams, and decisions
- interactive visualization for long-horizon project reconstruction
NeuralLog begins from a simple premise: high-value knowledge is frequently produced in chat, but chat systems are poor long-term memory systems. Recovering that knowledge usually requires either individual recollection or time-consuming manual searching through conversational archives.
This project explores whether a lightweight retrieval layer can materially improve that situation. In that sense, NeuralLog is not just a software utility; it is also an applied investigation into how any community or organization might preserve conversational knowledge as a searchable asset.


