deepseek_ocr_coc

DeepSeek-OCR

DeepSeek-OCR reimagines optical character recognition as a context compression problem — treating visual documents not as images to scan, but as information to compress and decode through an LLM-centric vision encoder. It converts documents, PDFs, and images to clean markdown, extracts text with layout awareness, parses figures, and localizes specific elements by reference — all at ~2500 tokens per second on a single A100 with vLLM. Multiple resolution modes from 64 to 400+ vision tokens let you tune the quality-speed tradeoff for your use case.

MIT
Image-Text-to-Text
PyTorch
Transformers
Safetensors
by @AIOZAI
60
0

Last updated: 3 months ago


Sign in to see model files

orCreate an account