Model Constraints
Some models have rules about their input — a max audio length, a fixed image size, or "exactly one face per query". If you break one, you get a clear error. Here's what to watch for per family.
Audio Models (CLAP)
| Constraint | Value | Notes |
|---|---|---|
| Max duration | 10 minutes (600 seconds) | Longer clips will error with AudioTooLongError |
| Sample rate | 48 kHz | Audio is automatically resampled |
| Max samples | 28,800,000 | 48,000 Hz × 600 seconds |
| Chunk size | 10 seconds | Long audio is automatically split into overlapping chunks |
| Chunk overlap | 1 second | Provides temporal continuity between chunks |
| Embedding mode | OneToMany | Returns one embedding per chunk with temporal metadata |
| Preprocessing | Required | NoPreprocessing not supported - always use ModelPreprocessing |
Audio chunking behavior:
- Audio ≤10 seconds: Returns 1 embedding
- Audio >10 seconds: Automatically chunked into 10-second segments with 1-second overlap
- Each chunk embedding includes metadata:
chunk_start_sec,chunk_end_sec,chunk_duration_sec,total_chunks,audio_total_duration_sec
Face Models (Buffalo_L, SFace+YuNet)
| Constraint | Value | Notes |
|---|---|---|
| Input size | 640x640 px | Images are resized internally |
| Face alignment | 112x112 px | Standard ArcFace alignment |
| Embedding mode | OneToMany | Returns one embedding per detected face |
| Preprocessing | Required | NoPreprocessing not supported |
| Query constraint | Single face | Query images must contain exactly 1 face |
Cross-modal pairs
Certain text and media encoders share a dimension so you can search across modalities — see Cross-modal pairs on the Models page.