Skip to main content
Arka vision tasks को support करता है: photo description, blueprint analysis, और screen capture।

Images describe करें

Two-layer analysis: OCR exact text extract करता है; vision layout और colors describe करता है।
Backends (auto-selected): Gemini, Ollama (llava), vLLM।

Drawings और blueprints

Gemini vision के साथ floor plans, elevations, MEP schematics, और scanned contracts analyze करें:

Screen capture

10-second countdown, full-display capture, vision describe:

Videos describe करें

describe_video ffmpeg के साथ frames sample करता है और उन frames को configured vision backend के माध्यम से भेजता है। इसका उपयोग gameplay recordings, UI animation checks, demos, और people-location questions के लिए करें।
People-focused prompts के लिए, Arka vision model से visible people की पहचान करने और उनकी approximate screen positions को पहचानने के लिए कहता है। सामान्य describe/analyze prompts के लिए, यह subjects, actions, setting, text, framing, और visual issues का वर्णन करता है।