The Solar Energy Catalog system provides high-performance caching for 15-minute solar energy period calculations, eliminating the 30-second bottleneck that previously occurred on every page load.
- 30x faster loads: <1s cached vs 30s uncached
- Incremental updates: Only processes new data since last period
- Persistent storage: Automatic saving to Storj with backup strategy
- Session-first approach: Instant loads after initial calculation
- Purpose: Manages storage, retrieval, and updates of energy catalogs
- Storage:
{mac_address}.energy_catalog.parquetin Storj bucket - Backup:
backups/{mac_address}.energy_catalog_{timestamp}.parquet - Methods:
catalog_exists(): Check if catalog exists in storageload_catalog(): Load existing catalog from storagedetect_and_calculate_periods(): Generate new catalog from dataupdate_catalog_with_new_data(): Incremental updatessave_catalog(): Persist to storage with backup
- Vectorized operations: Uses
pandas.resample()instead of O(n²) iteration - Streamlit caching:
@st.cache_data()with 20-entry, 2-hour TTL - Timezone handling: Proper UTC→Pacific conversion
- Auto-loading: Loads catalog during app initialization
- Incremental updates: Updates with new archive data
- Auto-save: Saves when catalog >2 days old
- User feedback: Progress messages during operations
- Status display: Shows catalog metrics and date range
- Management controls: Regenerate and save buttons
- Details view: Expandable catalog details
# During app startup - automatic loading
if "energy_catalog" not in st.session_state:
catalog = EnergyCatalog(device_mac)
if catalog.catalog_exists():
periods_df = catalog.load_catalog()
# Update with any new data
updated_df = catalog.update_catalog_with_new_data(history_df, periods_df)
st.session_state["energy_catalog"] = updated_df
else:
# Generate new catalog
periods_df = catalog.detect_and_calculate_periods(history_df)
st.session_state["energy_catalog"] = periods_df# From diagnostics UI
catalog = EnergyCatalog(device_mac)
periods_df = catalog.detect_and_calculate_periods(history_df, auto_save=False)
catalog.save_catalog(periods_df)| Scenario | Time | Improvement |
|---|---|---|
| First load (no cache) | ~30s | Same |
| Cached loads | <1s | 30x faster |
| Incremental updates | <5s | 6x faster |
| Algorithm only | ~6s | 5x faster |
- Location:
lookoutbucket in Storj - Format: Parquet with snappy compression
- Naming:
{mac_address}.energy_catalog.parquet
- Location:
backups/folder in same bucket - Naming:
{mac_address}.energy_catalog_{timestamp}.parquet - Trigger: Automatic before overwriting existing catalog
- Threshold: >2 days since last save
- Trigger: During data pipeline loading
- Purpose: Ensure data durability without excessive I/O
period_start: datetime64[ns, America/Los_Angeles] # Start of 15min period
period_end: datetime64[ns, America/Los_Angeles] # End of 15min period
energy_kwh: float # Energy in kWh
updated_at: datetime # Last update timestamp
catalog_version: str # Version identifierdate: datetime64[ns, America/Los_Angeles] # TZ-aware datetime
dateutc: int64 # Milliseconds timestamp
solarradiation: float # W/m²Cause: Data doesn't have expected 'datetime' column Solution: Ensure data has 'date' column (TZ-aware datetime)
Cause: Catalog loading failed during startup Solution: Check diagnostics tab for error details, try manual regeneration
Cause: First-time catalog generation Solution: Expected behavior, subsequent loads will be cached
- Go to Diagnostics tab
- Click "🔄 Regenerate Energy Catalog"
- Wait for completion message
- Go to Diagnostics tab
- View "Solar Energy Catalog Management" section
- Check metrics and date range
- Refresh browser page
- Catalog will reload from storage
- Unit tests:
tests/test_energy_catalog.py(13 tests) - Integration tests: Solar UI and data pipeline tests
- Performance tests:
verify_energy_catalog.pyscript
- Core logic:
lookout/core/energy_catalog.py - Algorithm:
lookout/core/solar_energy_periods.py - Pipeline:
lookout/core/data_processing.py - UI:
lookout/ui/diagnostics.py
- Compression optimization: Further reduce storage size
- Memory optimization: Reduce memory footprint for large catalogs
- Parallel processing: Multi-threaded calculation for very large datasets
- Advanced caching: LRU eviction for memory-constrained environments