Skip to content

Latest commit

 

History

History
483 lines (401 loc) · 14.9 KB

File metadata and controls

483 lines (401 loc) · 14.9 KB

SynthoraAI OCR - Architecture Documentation

System Overview

The SynthoraAI OCR system is a production-ready, distributed architecture designed for high-performance optical character recognition at scale. It supports multiple OCR engines, intelligent caching, job queuing, and real-time progress updates.

High-Level Architecture

┌─────────────────────────────────────────────────────────────────┐
│                        Client Applications                       │
│  ┌─────────────┐  ┌─────────────┐  ┌─────────────────────────┐ │
│  │   Web UI    │  │  Mobile App │  │  External Integrations  │ │
│  │  (React)    │  │   (REST)    │  │    (SynthoraAI)        │ │
│  └─────────────┘  └─────────────┘  └─────────────────────────┘ │
└────────────┬──────────────┬────────────────────┬────────────────┘
             │              │                    │
             ▼              ▼                    ▼
┌─────────────────────────────────────────────────────────────────┐
│                       API Gateway / Load Balancer                │
│                    (Nginx / ALB / Ingress)                       │
└────────────┬──────────────┬────────────────────┬────────────────┘
             │              │                    │
   ┌─────────▼─────────┐    │     ┌──────────────▼──────────────┐
   │   WebSocket       │    │     │   RESTful API (Node.js)     │
   │   Server          │    │     │                             │
   │                   │    │     │  - Authentication           │
   │  - Real-time      │    │     │  - Rate Limiting            │
   │    Updates        │    │     │  - Request Validation       │
   │  - Job Progress   │    │     │  - Job Management           │
   └───────────────────┘    │     └──────────────┬──────────────┘
                            │                    │
                            │     ┌──────────────▼──────────────┐
                            │     │    Queue System (Bull)      │
                            │     │                             │
                            │     │  - Job Scheduling           │
                            │     │  - Priority Queuing         │
                            │     │  - Retry Logic              │
                            │     │  - Dead Letter Queue        │
                            │     └──────────────┬──────────────┘
                            │                    │
                            │     ┌──────────────▼──────────────┐
                            └────▶│   OCR Processing Backend    │
                                  │      (Python/FastAPI)       │
                                  │                             │
                                  │  ┌─────────────────────┐   │
                                  │  │  Image Preprocessor │   │
                                  │  └──────────┬──────────┘   │
                                  │             │               │
                                  │  ┌──────────▼──────────┐   │
                                  │  │   OCR Engines       │   │
                                  │  │  - Tesseract        │   │
                                  │  │  - EasyOCR          │   │
                                  │  │  - PaddleOCR        │   │
                                  │  │  - TrOCR            │   │
                                  │  │  - Ensemble         │   │
                                  │  └──────────┬──────────┘   │
                                  │             │               │
                                  │  ┌──────────▼──────────┐   │
                                  │  │  Post-Processor     │   │
                                  │  │  - Text Correction  │   │
                                  │  │  - Confidence Calc  │   │
                                  │  └─────────────────────┘   │
                                  └──────────────┬──────────────┘
                                                 │
                ┌────────────────────────────────┼────────────────────────────┐
                │                                │                            │
    ┌───────────▼──────────┐       ┌────────────▼───────────┐   ┌────────────▼──────────┐
    │   Cache Layer        │       │   Database Layer       │   │   Storage Layer       │
    │   (Redis)            │       │   (MongoDB)            │   │   (S3/Blob)          │
    │                      │       │                        │   │                       │
    │  - Result Caching    │       │  - Job Metadata        │   │  - Original Files     │
    │  - Session Storage   │       │  - User Data           │   │  - Processed Results  │
    │  - Rate Limit State  │       │  - Processing History  │   │  - Model Artifacts    │
    └──────────────────────┘       └────────────────────────┘   └───────────────────────┘

Component Details

1. API Layer (Node.js/Express)

Responsibilities:

  • HTTP request handling and routing
  • Authentication and authorization (JWT/API keys)
  • Rate limiting and throttling
  • Input validation and sanitization
  • WebSocket management for real-time updates
  • Job queue management

Key Files:

  • ocr_api/index.js - Main application entry point
  • ocr_api/routes/ocr.js - OCR endpoint definitions
  • ocr_api/controllers/ocrController.js - Request handling logic
  • ocr_api/middleware/auth.js - Authentication middleware
  • ocr_api/middleware/rateLimiter.js - Rate limiting logic

Scalability:

  • Stateless design for horizontal scaling
  • Load balancer compatible
  • Supports multiple instances

2. OCR Processing Backend (Python/FastAPI)

Responsibilities:

  • OCR engine orchestration
  • Image preprocessing
  • Multi-engine processing
  • Result aggregation (ensemble mode)
  • Performance optimization

Key Components:

Image Preprocessor

class ImagePreprocessor:
    - Grayscale conversion
    - Noise removal (denoising)
    - Skew correction (deskewing)
    - Contrast enhancement (CLAHE)
    - Binarization (Otsu's method)
    - Border removal
    - Resolution optimization

OCR Processor

class OCRProcessor:
    - Engine initialization
    - Image processing pipeline
    - Confidence scoring
    - Language detection
    - Ensemble voting

PDF Processor

class PDFProcessor:
    - Text extraction (searchable PDFs)
    - Image conversion (scanned PDFs)
    - Page-by-page processing
    - Embedded image extraction

Supported Engines:

  1. Tesseract - Fast, supports 100+ languages
  2. EasyOCR - Deep learning-based, high accuracy
  3. PaddleOCR - Ultra-fast, good balance
  4. TrOCR - Transformer-based, best for handwriting
  5. Ensemble - Combines multiple engines for best results

3. Job Queue System (Bull/Redis)

Responsibilities:

  • Asynchronous job processing
  • Priority queue management
  • Job retry with exponential backoff
  • Failed job handling
  • Progress tracking

Queue Types:

ocr-processing     → Standard OCR jobs
batch-processing   → Batch operations (5 concurrent)
priority-processing → High-priority jobs

Job States:

pending → processing → completed
                   ↓
                failed → retry (up to 3 times) → dead-letter

4. WebSocket Service

Responsibilities:

  • Real-time client connections
  • Job progress updates
  • Event broadcasting
  • Subscription management

Events:

// Client → Server
subscribe(jobId)
unsubscribe(jobId)
ping()

// Server → Client
connected
progress(jobId, percent, message)
complete(jobId, result)
error(jobId, error)

5. Database Layer (MongoDB)

Collections:

Jobs Collection

{
  jobId: String,
  userId: String,
  type: 'image' | 'pdf' | 'batch' | 'url',
  status: 'pending' | 'processing' | 'completed' | 'failed',
  engine: String,
  input: { filename, size, mimetype, url },
  result: { text, confidence, processingTime, pages, metadata },
  error: { message, stack, code },
  priority: Number,
  retries: Number,
  createdAt: Date,
  startedAt: Date,
  completedAt: Date,
  expiresAt: Date
}

Indexes:

  • jobId (unique)
  • userId + createdAt
  • status + priority + createdAt
  • expiresAt (TTL index, auto-delete after 7 days)

6. Cache Layer (Redis)

Use Cases:

  1. Result Caching

    • Key: ocr:{engine}:{languages}:{image_hash}
    • TTL: 1 hour
    • Saves redundant processing
  2. Rate Limiting

    • Sliding window algorithm
    • Per-IP or per-user limits
    • Different tiers (free, premium)
  3. Session Storage

    • User sessions
    • Temporary job data
    • WebSocket subscriptions
  4. Job Queue State

    • Bull queue storage
    • Job progress tracking

7. Storage Layer

Cloud Storage Integration:

  • AWS S3 - Primary storage
  • Azure Blob - Alternative
  • Google Cloud Storage - Alternative

Storage Structure:

/uploads/{userId}/{timestamp}/original.{ext}
/results/{userId}/{timestamp}/result.json
/results/{userId}/{timestamp}/processed.{ext}
/models/{engine}/{version}/model.{ext}

Data Flow

1. Synchronous Processing

Client Request
    ↓
API Validation
    ↓
Authentication/Rate Limiting
    ↓
Cache Check (hit → return cached result)
    ↓
Submit to Python Backend
    ↓
Image Preprocessing
    ↓
OCR Engine Processing
    ↓
Post-Processing
    ↓
Cache Result
    ↓
Return to Client

2. Asynchronous Processing (Queue-based)

Client Request
    ↓
API Validation
    ↓
Create Job Record (MongoDB)
    ↓
Add to Queue (Bull/Redis)
    ↓
Return Job ID to Client
    ↓
Client Subscribes (WebSocket)
    ↓
Worker Picks Job from Queue
    ↓
Processing (with progress updates)
    ↓
Update Job Record
    ↓
Notify Client (WebSocket)
    ↓
Client Retrieves Result

Security Architecture

Authentication Layers

  1. API Key Authentication

    Header: X-API-Key: your-api-key
    
  2. JWT Authentication

    Header: Authorization: Bearer {token}
    
  3. OAuth 2.0 (Optional)

    • Integration with external providers
    • Social login support

Security Measures

  1. Input Validation

    • File type checking
    • Size limits
    • Malware scanning (optional)
  2. Rate Limiting

    • IP-based limits
    • User-based limits
    • Tier-based quotas
  3. Data Encryption

    • HTTPS/TLS for transit
    • AES-256 for storage
    • Database encryption at rest
  4. Access Control

    • Role-based access control (RBAC)
    • Resource ownership validation
    • Audit logging

Scalability & Performance

Horizontal Scaling

API Layer:

  • Stateless design
  • Load balancer distribution
  • Auto-scaling based on CPU/memory

Processing Backend:

  • Worker pool architecture
  • GPU acceleration support
  • Parallel processing for batch jobs

Caching Strategy

Levels:

  1. L1: In-Memory (Process-level)
  2. L2: Redis (Distributed)
  3. L3: CDN (Edge caching)

Cache Invalidation:

  • TTL-based expiration
  • Manual invalidation API
  • Pattern-based clearing

Performance Optimizations

  1. Image Preprocessing

    • Lazy loading
    • Async processing
    • Format conversion
  2. OCR Processing

    • GPU acceleration (CUDA)
    • Model quantization
    • Batch processing
  3. Network

    • Response compression (gzip)
    • HTTP/2 support
    • Connection pooling

Monitoring & Observability

Metrics (Prometheus)

# Request metrics
ocr_requests_total{engine, status}
ocr_processing_duration_seconds{engine}
ocr_confidence_score{engine}
ocr_errors_total{engine, error_type}
ocr_active_requests
ocr_queue_size

# System metrics
system_cpu_usage
system_memory_usage
system_disk_usage

Logging

Structured JSON Logs:

{
  "timestamp": "2025-01-15T10:30:00Z",
  "level": "info",
  "service": "ocr-backend",
  "message": "Image processed successfully",
  "metadata": {
    "job_id": "abc123",
    "engine": "tesseract",
    "processing_time": 1.23,
    "confidence": 0.95
  }
}

Distributed Tracing

  • OpenTelemetry integration
  • Request correlation IDs
  • Service dependency mapping

Disaster Recovery

Backup Strategy

  1. Database Backups

    • Daily automated backups
    • Point-in-time recovery
    • Cross-region replication
  2. File Storage Backups

    • Versioning enabled
    • Lifecycle policies
    • Disaster recovery region

High Availability

  • Multi-region deployment
  • Automated failover
  • Health check monitoring
  • Circuit breaker pattern

Cost Optimization

  1. Compute

    • Spot instances for batch processing
    • Auto-scaling to match demand
    • Serverless for sporadic workloads
  2. Storage

    • Lifecycle policies (archive old data)
    • Compression
    • Tiered storage
  3. Network

    • CDN for static assets
    • Data transfer optimization
    • Regional optimization

Last Updated: 2025-01-15 Version: 1.0.0 Maintainer: SynthoraAI Team