AI Workers (oridecon-ai-workers)
AI background workers for the Oridecon Framework — batch embedding, document ingestion, DLQ, maintenance
Overview
Section titled “Overview”AI background workers for the Oridecon Framework. Handles the heavy-lifting off the request path: document ingestion, batch embedding generation, periodic maintenance, and dead-letter-queue recovery — all with progress tracking, exponential backoff, and health reporting. Zero-config usage starts with sensible defaults.
Full documentation: docs.oridecon.dev
Install
Section titled “Install”uv add oridecon-ai-workersQuick Start
Section titled “Quick Start”from oridecon import Applicationfrom oridecon.di.module import Module, module
from oridecon.ai.workers import WorkersModulefrom oridecon.ai.workers.config import WorkersConfig
@module( imports=[ WorkersModule.configure( WorkersConfig( batch_embedding_concurrency=3, document_ingestion_concurrency=3, enable_maintenance=True, dlq_check_interval=60, ) ) ])class AppModule(Module): pass
async with Application.boot(modules=[AppModule]) as app: # use app.container to resolve services ...Configuration
Section titled “Configuration”Zero-config usage: Call
WorkersModule.configure()with no arguments to use defaults.
Option 1 — YAML file
Section titled “Option 1 — YAML file”ai_workers: enabled: true batch_embedding_concurrency: 3 document_ingestion_concurrency: 3 enable_maintenance: true dlq_check_interval: 60Option 2 — Profiles + Environment Variables (recommended)
Section titled “Option 2 — Profiles + Environment Variables (recommended)”export ORI_AI_WORKERS__BATCH_EMBEDDING_CONCURRENCY=5# Environment variables for each fieldOption 3 — Python
Section titled “Option 3 — Python”from oridecon.ai.workers.config import WorkersConfigfrom oridecon.ai.workers import WorkersModule
config = WorkersConfig( enabled=True, batch_embedding_concurrency=5, document_ingestion_concurrency=3, enable_maintenance=True, dlq_check_interval=60,)WorkersModule.configure(config)Config reference
Section titled “Config reference”| Field | Default | Env var | Description |
|---|---|---|---|
enabled | True | ORI_AI_WORKERS__ENABLED | Master on/off switch for all background workers |
batch_embedding_concurrency | 3 | ORI_AI_WORKERS__BATCH_EMBEDDING_CONCURRENCY | Concurrent embedding batch tasks |
document_ingestion_concurrency | 3 | ORI_AI_WORKERS__DOCUMENT_INGESTION_CONCURRENCY | Concurrent document processing tasks |
enable_maintenance | True | ORI_AI_WORKERS__ENABLE_MAINTENANCE | Enable vector-store and cache maintenance |
dlq_check_interval | 60 | ORI_AI_WORKERS__DLQ_CHECK_INTERVAL | Seconds between DLQ recovery sweeps |
Securing document ingestion (path-traversal control)
Section titled “Securing document ingestion (path-traversal control)”When a document source path can be influenced by users (an upload filename,
an API parameter), constrain it with allowed_root. Both the worker and the
underlying parser accept the same opt-in argument:
from pathlib import Pathfrom oridecon.ai.workers.document_ingestion import DocumentIngestionWorker
worker = DocumentIngestionWorker( vector_store=store, queue=queue, allowed_root=Path("/srv/app/documents"),)With allowed_root set, every ingested source is resolved — symlinks and
.. segments followed — and must land inside that directory, otherwise the
job fails with RAGError. The default (allowed_root=None) performs no
containment and preserves historical behavior, which is appropriate when all
sources are fully trusted server-local files. A custom document_parser
passed to the worker is responsible for its own path policy; allowed_root
only applies to the built-in UniversalDocumentParser the worker constructs.
The same check applies to UniversalDocumentParser.parse() and
UniversalDocumentParser.extract_metadata() when the parser is used directly,
and to LoaderWorkerBridge sources (the bridge submits to the worker, so
configure allowed_root on the worker).
Module Factory Methods
Section titled “Module Factory Methods”| Method | Description |
|---|---|
WorkersModule.configure(config, enable_scheduler) | Configure with explicit config |
WorkersModule.stub(config) | Minimal config for testing |
Key Features
Section titled “Key Features”- Document ingestion worker: Parse PDF, DOCX, TXT, HTML, Markdown into chunks for vector store
- Batch embedding worker: Process chunks in configurable batches with in-memory embedding cache
- Dead letter queue worker: Handle failed jobs with failure classification and exponential backoff
- Maintenance worker: Periodic vector store index optimization, cache cleanup, document cleanup
- Progress tracking: Job progress monitoring with cache hit rate statistics
- Adapters: RAGAdapter, TasksAdapter, LoaderWorker for ecosystem integration
Testing
Section titled “Testing”async with Application.boot(modules=[WorkersModule.stub()]) as app: # your test code ...Key Source Files
Section titled “Key Source Files”| File | What it contains |
|---|---|
src/oridecon/ai/workers/module.py | WorkersModule.configure(), .stub() |
src/oridecon/ai/workers/config.py | WorkersConfig |
src/oridecon/ai/workers/document_ingestion/worker.py | DocumentIngestionWorker |
src/oridecon/ai/workers/batch_embedding/worker.py | BatchEmbeddingWorker |
src/oridecon/ai/workers/dlq/worker.py | DeadLetterQueueWorker |
src/oridecon/ai/workers/maintenance/worker.py | MaintenanceWorker |
src/oridecon/ai/workers/types.py | DLQItem, DLQStats, MaintenanceTask, MaintenanceResult |
src/oridecon/ai/workers/di/provider.py | WorkersProvider |