The SynthoraAI OCR system is a production-ready, distributed architecture designed for high-performance optical character recognition at scale. It supports multiple OCR engines, intelligent caching, job queuing, and real-time progress updates.
┌─────────────────────────────────────────────────────────────────┐
│ Client Applications │
│ ┌─────────────┐ ┌─────────────┐ ┌─────────────────────────┐ │
│ │ Web UI │ │ Mobile App │ │ External Integrations │ │
│ │ (React) │ │ (REST) │ │ (SynthoraAI) │ │
│ └─────────────┘ └─────────────┘ └─────────────────────────┘ │
└────────────┬──────────────┬────────────────────┬────────────────┘
│ │ │
▼ ▼ ▼
┌─────────────────────────────────────────────────────────────────┐
│ API Gateway / Load Balancer │
│ (Nginx / ALB / Ingress) │
└────────────┬──────────────┬────────────────────┬────────────────┘
│ │ │
┌─────────▼─────────┐ │ ┌──────────────▼──────────────┐
│ WebSocket │ │ │ RESTful API (Node.js) │
│ Server │ │ │ │
│ │ │ │ - Authentication │
│ - Real-time │ │ │ - Rate Limiting │
│ Updates │ │ │ - Request Validation │
│ - Job Progress │ │ │ - Job Management │
└───────────────────┘ │ └──────────────┬──────────────┘
│ │
│ ┌──────────────▼──────────────┐
│ │ Queue System (Bull) │
│ │ │
│ │ - Job Scheduling │
│ │ - Priority Queuing │
│ │ - Retry Logic │
│ │ - Dead Letter Queue │
│ └──────────────┬──────────────┘
│ │
│ ┌──────────────▼──────────────┐
└────▶│ OCR Processing Backend │
│ (Python/FastAPI) │
│ │
│ ┌─────────────────────┐ │
│ │ Image Preprocessor │ │
│ └──────────┬──────────┘ │
│ │ │
│ ┌──────────▼──────────┐ │
│ │ OCR Engines │ │
│ │ - Tesseract │ │
│ │ - EasyOCR │ │
│ │ - PaddleOCR │ │
│ │ - TrOCR │ │
│ │ - Ensemble │ │
│ └──────────┬──────────┘ │
│ │ │
│ ┌──────────▼──────────┐ │
│ │ Post-Processor │ │
│ │ - Text Correction │ │
│ │ - Confidence Calc │ │
│ └─────────────────────┘ │
└──────────────┬──────────────┘
│
┌────────────────────────────────┼────────────────────────────┐
│ │ │
┌───────────▼──────────┐ ┌────────────▼───────────┐ ┌────────────▼──────────┐
│ Cache Layer │ │ Database Layer │ │ Storage Layer │
│ (Redis) │ │ (MongoDB) │ │ (S3/Blob) │
│ │ │ │ │ │
│ - Result Caching │ │ - Job Metadata │ │ - Original Files │
│ - Session Storage │ │ - User Data │ │ - Processed Results │
│ - Rate Limit State │ │ - Processing History │ │ - Model Artifacts │
└──────────────────────┘ └────────────────────────┘ └───────────────────────┘
Responsibilities:
- HTTP request handling and routing
- Authentication and authorization (JWT/API keys)
- Rate limiting and throttling
- Input validation and sanitization
- WebSocket management for real-time updates
- Job queue management
Key Files:
ocr_api/index.js- Main application entry pointocr_api/routes/ocr.js- OCR endpoint definitionsocr_api/controllers/ocrController.js- Request handling logicocr_api/middleware/auth.js- Authentication middlewareocr_api/middleware/rateLimiter.js- Rate limiting logic
Scalability:
- Stateless design for horizontal scaling
- Load balancer compatible
- Supports multiple instances
Responsibilities:
- OCR engine orchestration
- Image preprocessing
- Multi-engine processing
- Result aggregation (ensemble mode)
- Performance optimization
Key Components:
class ImagePreprocessor:
- Grayscale conversion
- Noise removal (denoising)
- Skew correction (deskewing)
- Contrast enhancement (CLAHE)
- Binarization (Otsu's method)
- Border removal
- Resolution optimizationclass OCRProcessor:
- Engine initialization
- Image processing pipeline
- Confidence scoring
- Language detection
- Ensemble votingclass PDFProcessor:
- Text extraction (searchable PDFs)
- Image conversion (scanned PDFs)
- Page-by-page processing
- Embedded image extractionSupported Engines:
- Tesseract - Fast, supports 100+ languages
- EasyOCR - Deep learning-based, high accuracy
- PaddleOCR - Ultra-fast, good balance
- TrOCR - Transformer-based, best for handwriting
- Ensemble - Combines multiple engines for best results
Responsibilities:
- Asynchronous job processing
- Priority queue management
- Job retry with exponential backoff
- Failed job handling
- Progress tracking
Queue Types:
ocr-processing → Standard OCR jobs
batch-processing → Batch operations (5 concurrent)
priority-processing → High-priority jobs
Job States:
pending → processing → completed
↓
failed → retry (up to 3 times) → dead-letter
Responsibilities:
- Real-time client connections
- Job progress updates
- Event broadcasting
- Subscription management
Events:
// Client → Server
subscribe(jobId)
unsubscribe(jobId)
ping()
// Server → Client
connected
progress(jobId, percent, message)
complete(jobId, result)
error(jobId, error)Collections:
{
jobId: String,
userId: String,
type: 'image' | 'pdf' | 'batch' | 'url',
status: 'pending' | 'processing' | 'completed' | 'failed',
engine: String,
input: { filename, size, mimetype, url },
result: { text, confidence, processingTime, pages, metadata },
error: { message, stack, code },
priority: Number,
retries: Number,
createdAt: Date,
startedAt: Date,
completedAt: Date,
expiresAt: Date
}Indexes:
jobId(unique)userId + createdAtstatus + priority + createdAtexpiresAt(TTL index, auto-delete after 7 days)
Use Cases:
-
Result Caching
- Key:
ocr:{engine}:{languages}:{image_hash} - TTL: 1 hour
- Saves redundant processing
- Key:
-
Rate Limiting
- Sliding window algorithm
- Per-IP or per-user limits
- Different tiers (free, premium)
-
Session Storage
- User sessions
- Temporary job data
- WebSocket subscriptions
-
Job Queue State
- Bull queue storage
- Job progress tracking
Cloud Storage Integration:
- AWS S3 - Primary storage
- Azure Blob - Alternative
- Google Cloud Storage - Alternative
Storage Structure:
/uploads/{userId}/{timestamp}/original.{ext}
/results/{userId}/{timestamp}/result.json
/results/{userId}/{timestamp}/processed.{ext}
/models/{engine}/{version}/model.{ext}
Client Request
↓
API Validation
↓
Authentication/Rate Limiting
↓
Cache Check (hit → return cached result)
↓
Submit to Python Backend
↓
Image Preprocessing
↓
OCR Engine Processing
↓
Post-Processing
↓
Cache Result
↓
Return to Client
Client Request
↓
API Validation
↓
Create Job Record (MongoDB)
↓
Add to Queue (Bull/Redis)
↓
Return Job ID to Client
↓
Client Subscribes (WebSocket)
↓
Worker Picks Job from Queue
↓
Processing (with progress updates)
↓
Update Job Record
↓
Notify Client (WebSocket)
↓
Client Retrieves Result
-
API Key Authentication
Header: X-API-Key: your-api-key -
JWT Authentication
Header: Authorization: Bearer {token} -
OAuth 2.0 (Optional)
- Integration with external providers
- Social login support
-
Input Validation
- File type checking
- Size limits
- Malware scanning (optional)
-
Rate Limiting
- IP-based limits
- User-based limits
- Tier-based quotas
-
Data Encryption
- HTTPS/TLS for transit
- AES-256 for storage
- Database encryption at rest
-
Access Control
- Role-based access control (RBAC)
- Resource ownership validation
- Audit logging
API Layer:
- Stateless design
- Load balancer distribution
- Auto-scaling based on CPU/memory
Processing Backend:
- Worker pool architecture
- GPU acceleration support
- Parallel processing for batch jobs
Levels:
- L1: In-Memory (Process-level)
- L2: Redis (Distributed)
- L3: CDN (Edge caching)
Cache Invalidation:
- TTL-based expiration
- Manual invalidation API
- Pattern-based clearing
-
Image Preprocessing
- Lazy loading
- Async processing
- Format conversion
-
OCR Processing
- GPU acceleration (CUDA)
- Model quantization
- Batch processing
-
Network
- Response compression (gzip)
- HTTP/2 support
- Connection pooling
# Request metrics
ocr_requests_total{engine, status}
ocr_processing_duration_seconds{engine}
ocr_confidence_score{engine}
ocr_errors_total{engine, error_type}
ocr_active_requests
ocr_queue_size
# System metrics
system_cpu_usage
system_memory_usage
system_disk_usage
Structured JSON Logs:
{
"timestamp": "2025-01-15T10:30:00Z",
"level": "info",
"service": "ocr-backend",
"message": "Image processed successfully",
"metadata": {
"job_id": "abc123",
"engine": "tesseract",
"processing_time": 1.23,
"confidence": 0.95
}
}- OpenTelemetry integration
- Request correlation IDs
- Service dependency mapping
-
Database Backups
- Daily automated backups
- Point-in-time recovery
- Cross-region replication
-
File Storage Backups
- Versioning enabled
- Lifecycle policies
- Disaster recovery region
- Multi-region deployment
- Automated failover
- Health check monitoring
- Circuit breaker pattern
-
Compute
- Spot instances for batch processing
- Auto-scaling to match demand
- Serverless for sporadic workloads
-
Storage
- Lifecycle policies (archive old data)
- Compression
- Tiered storage
-
Network
- CDN for static assets
- Data transfer optimization
- Regional optimization
Last Updated: 2025-01-15 Version: 1.0.0 Maintainer: SynthoraAI Team