Skip to content

Latest commit

Β 

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Vibe Coding a Rust ML Pipeline

A workshop demo project that builds an Iris flower classifier in Rust using Linfa, demonstrating the "vibe coding" workflow -- where AI assistants handle syntax while you focus on intent, and the Rust compiler acts as your safety net.

What It Does

flowchart LR
    A["πŸ“¦ Load Dataset"] --> B["βœ‚οΈ Train/Test Split"]
    B --> C["🧠 Train Models"]
    C --> D["πŸ“Š Evaluate"]
    D --> E["πŸ–₯️ Terminal Output"]
    style A fill:#e1f5fe
    style C fill:#fff3e0
    style D fill:#e8f5e9
Loading
  • Loads the classic Iris dataset (150 samples, 4 features, 3 classes)
  • Trains two Decision Tree classifiers (participant model from lib.rs + built-in Entropy tree)
  • Evaluates both models and displays results in formatted terminal tables

Prerequisites

  • Rust toolchain (1.93+ stable): Install via rustup
  • Git (to navigate workshop checkpoints)

Quick Start

git clone https://github.com/Tony363/vibe-rust-ml-workshop.git
cd vibe-rust-ml-workshop
cargo run --release

Workshop Steps

The project is built incrementally. Each Git tag represents a compilable, runnable checkpoint:

flowchart LR
    S1["step-1-scaffold\nπŸ—οΈ Project Setup"]
    S2["step-2-data\nπŸ“¦ Data Loading"]
    S3["step-3-training\n🧠 Model Training"]
    S4["step-4-complete\nπŸ“Š Evaluation"]
    S1 --> S2 --> S3 --> S4
    style S1 fill:#e3f2fd
    style S2 fill:#e8f5e9
    style S3 fill:#fff3e0
    style S4 fill:#fce4ec
Loading

Step 1: Scaffold (step-1-scaffold)

git checkout step-1-scaffold
cargo run
# Output: "Vibe Rust ML Workshop"

Minimal project setup with dependencies in Cargo.toml: linfa, linfa-trees, linfa-datasets, ndarray, comfy-table, rand.

Step 2: Data Loading (step-2-data)

git checkout step-2-data
cargo run

Loads Iris dataset via linfa_datasets::iris(), displays a dataset info table, and splits 80/20 for training/testing.

Step 3: Model Training (step-3-training)

git checkout step-3-training
cargo run

Trains two Decision Tree classifiers:

  • Tree 1 (lib.rs): Entropy, max depth = 10
  • Tree 2 (main.rs): Entropy, unlimited depth

Step 4: Evaluation + Pretty Output (step-4-complete)

git checkout step-4-complete
cargo run --release

Full pipeline with predictions, accuracy comparison, confusion matrix, and sample predictions table.

To return to the final version:

git checkout master

Expected Output

  Vibe Rust ML Workshop -- Iris Classification

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Property      ┆ Value                                                β”‚
β•žβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•ͺ══════════════════════════════════════════════════════║
β”‚ Dataset       ┆ Iris                                                 β”‚
β”œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”€
β”‚ Samples       ┆ 150                                                  β”‚
β”œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”€
β”‚ Features      ┆ 4                                                    β”‚
β”œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”€
β”‚ Classes       ┆ 3 (Setosa, Versicolor, Virginica)                    β”‚
β”œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”€
β”‚ Feature Names ┆ sepal length, sepal width, petal length, petal width β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

  Train/Test split: 120 training, 30 testing samples

  Training Model 1: DecisionTree (Entropy, depth=10)...
  -> Trained in ~130Β΅s
  Training Model 2: Decision Tree (Entropy, unlimited depth)...
  -> Trained in ~150Β΅s

  Model Comparison
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Model                            ┆ Split Quality ┆ Max Depth ┆ Accuracy ┆ Train Time β”‚
β•žβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•ͺ═══════════════β•ͺ═══════════β•ͺ══════════β•ͺ════════════║
β”‚ DecisionTree (Entropy, depth=10) ┆ -             ┆ -         ┆ ~93-100% ┆ ~130Β΅s     β”‚
β”œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”€
β”‚ Tree 2                           ┆ Entropy       ┆ None      ┆ ~93-100% ┆ ~150Β΅s     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

  Confusion Matrix (DecisionTree (Entropy, depth=10))
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Actual \ Predicted ┆ Setosa ┆ Versicolor ┆ Virginica β”‚
β•žβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•ͺ════════β•ͺ════════════β•ͺ═══════════║
β”‚ Setosa             ┆ ~10    ┆ 0          ┆ 0         β”‚
β”œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”€
β”‚ Versicolor         ┆ 0      ┆ ~10        ┆ 0-1       β”‚
β”œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”€
β”‚ Virginica          ┆ 0      ┆ 0-1        ┆ ~10       β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

  Sample Predictions (first 10 test samples)
β”Œβ”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ #  ┆ Sepal L ┆ Sepal W ┆ Petal L ┆ Petal W ┆ Actual     ┆ Predicted  β”‚
β•žβ•β•β•β•β•ͺ═════════β•ͺ═════════β•ͺ═════════β•ͺ═════════β•ͺ════════════β•ͺ════════════║
β”‚ 1  ┆ 5.0     ┆ 3.4     ┆ 1.5     ┆ 0.2     ┆ Setosa     ┆ Setosa     β”‚
β”œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”€
β”‚ 2  ┆ 6.8     ┆ 3.2     ┆ 5.9     ┆ 2.3     ┆ Virginica  ┆ Virginica  β”‚
β”œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”€
β”‚ .. ┆ ...     ┆ ...     ┆ ...     ┆ ...     ┆ ...        ┆ ...        β”‚
β”œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”Όβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ•Œβ”€
β”‚ 10 ┆ 6.9     ┆ 3.1     ┆ 5.4     ┆ 2.1     ┆ Virginica  ┆ Virginica  β”‚
β””β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Note: the demo uses seed 42; the CI scorer (src/bin/score.rs) uses seed 1 for fair leaderboard comparison.

Project Structure

graph LR
    subgraph "Source Code"
        lib["src/lib.rs\n✏️ Participant Models"]
        main["src/main.rs\nDemo Binary"]
        score["src/bin/score.rs\n🌸 Iris Scorer"]
        wine["src/bin/score_wine.rs\n🍷 Wine Scorer"]
    end
    subgraph "CI / Leaderboard"
        lb["leaderboard.yml\nScore PRs"]
        ulb["update-leaderboard.yml\nUpdate Rankings"]
        cs["check-steps.yml\nVerify Tags"]
    end
    subgraph "Docs"
        ws["WORKSHOP.md"]
        rules["LEADERBOARD_RULES.md"]
        lead["LEADERBOARD.md"]
    end
    lib --> main
    lib --> score
    lib --> wine
    score --> lb
    wine --> lb
    lb --> lead
    ulb --> lead
Loading

Dependencies

Crate Purpose
linfa ML framework (scikit-learn for Rust)
linfa-trees Decision Tree classifier
linfa-datasets Built-in datasets (Iris, Wine Quality)
ndarray N-dimensional arrays
comfy-table Pretty terminal tables
rand Random number generation (shuffling)

Workshop Challenge

Think you can beat the baseline? Two challenges, pick one or both!

xychart-beta horizontal
    title "Baseline Accuracy β€” Can You Beat It?"
    x-axis ["🌸 Iris (Easy)", "🍷 Wine Quality (Hard)"]
    bar [93.3, 53.9]
Loading
Challenge Dataset Baseline Beat this!
Iris (Easy) 150 samples, 4 features, 3 classes 93.3% Can you hit 100%?
Wine Quality (Hard) 1599 samples, 11 features, 6 classes 53.9% Can you break 65%?

Quick Start

git checkout submissions
git checkout -b my-submission
# edit src/lib.rs -- change build_and_predict() and/or build_and_predict_wine()
cargo run --bin score --release       # check Iris score locally
cargo run --bin score_wine --release  # check Wine Quality score locally
git add -A && git commit -m "my submission"
git push -u origin my-submission
# open a PR targeting the 'submissions' branch

Rules

  • Only modify src/lib.rs and Cargo.toml
  • Must use linfa algorithms
  • CI scores both challenges with a fixed seed for fairness
  • CI will automatically post a combined leaderboard on your PR

See LEADERBOARD.md for current standings and LEADERBOARD_RULES.md for full rules.

Resources

License

MIT

About

Workshop demo: Vibe coding a Rust ML pipeline with Linfa (Iris classification, Decision Trees, pretty CLI output)

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages