File: backend/app/graph/nodes/input_ingest.py.
Entry point of the graph. Loads the CSV via DuckDB, detects key columns heuristically, and packages input_context + input_constraints for downstream nodes.
| Key | Used for |
|---|---|
raw_csv_path |
File to load. May be a basename — resolved against /tmp/retain_ai_uploads/ first, then os.getcwd()/data/. |
questionnaire |
Pull-through to input_context (business_context, industry, company_size) and input_constraints (time_range, product_lines, market_segment, budget, legal_constraints). |
df = duckdb.connect(":memory:").execute(f"SELECT * FROM read_csv_auto('{path}')").df()
# Case-insensitive heuristic column detection
cols_lower = {col.lower(): col for col in df.columns}
customer_id_col = next((cols_lower[k] for k in cols_lower if 'id' in k or 'user' in k), None)
tenure_col = next((cols_lower[k] for k in cols_lower if 'tenure' in k or 'months_active' in k or k == 'months'), None)
usage_col = next((cols_lower[k] for k in cols_lower if 'usage' in k or 'logins' in k), None)
support_col = next((cols_lower[k] for k in cols_lower if 'support' in k or 'tickets' in k), None)
plan_col = next((cols_lower[k] for k in cols_lower if 'plan' in k), None) \
or next((cols_lower[k] for k in cols_lower if 'contract' in k), None)
contract_col = next((cols_lower[k] for k in cols_lower if 'contract' in k and cols_lower[k] != plan_col), None)
churn_col = get_churn_column(df) # in app/graph/utils.pyget_churn_column() is stricter: dtype must be int/float AND values must be a subset of {0, 1}. Falls back to columns literally named is_churned / churned.
{
"raw_csv_path": <resolved absolute path>, # rewritten so downstream nodes don't depend on cwd
"normalized_df": [{...row dict...}, ...], # whole CSV
"input_context": {
"source": path,
"row_count": int,
"column_count": int,
"detected_columns": {
"customer_id", "tenure", "usage", "support", "plan", "contract", "churn" # values or None
},
"business_context": str,
"industry": str,
"company_size": str,
},
"input_constraints": {
"time_range": str,
"product_lines": list,
"market_segment": str,
"budget_constraints": str,
"legal_constraints": list,
},
"current_node": "input_ingest",
"retry_count": int,
}read_csv_auto figures out delimiters, header rows, and dtypes without configuration. Cheaper than Pandas' inference for the 50–10k-row datasets the pipeline targets.
On any exception (file not found, malformed CSV, encoding error) the node returns:
{
"errors": [*state.get("errors", []), f"Input ingest error: {e}"],
"current_node": "input_ingest",
"retry_count": state.get("retry_count", 0) + 1,
}The graph continues with a missing normalized_df. data_audit will then fail similarly and route_after_data_audit will route to retry_handler (which currently exits immediately since MAX_RETRIES=0).
- Graph entry point (
graph.set_entry_point("input_ingest")). - Re-entered from
retry_handlerifretry_count < MAX_RETRIES.
normalized_dfis the entire CSV as a list of row dicts. For a 10k-row file that's measurable RSS pressure on Render's free tier. Downstream nodes re-read the CSV fromraw_csv_pathvia DuckDB instead of pulling from state — this state key exists for completeness but is essentially write-only.detected_columnsvalues may beNone. Every downstream node guards withnext((c for c in df.columns if ...), None)as a fallback heuristic in case ingest missed.planandcontractare detected as separate columns (e.g. "Plan Tier" vs "Contract Length") —planprefers a literalplanmatch first, falling back tocontractonly if no plan column exists;contractthen looks for a second, distinctcontract-matching column. This split feeds the forensicchurn_by_contractstat bucket and the CoxPH one-hot encoding — contract cadence is often the single strongest churn split in subscription data and was previously merged intoplan, hiding it.- The case-insensitive scan is greedy and order-dependent — if your CSV has both
customer_idanduser_idcolumns, the first match wins (dict iteration order).