Lightweight Professional OCR Model
GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder-decoder architecture. The model integrates the CogViT visual encoder pre-trained on large-scale image-text data, a lightweight cross-modal connector with efficient token downsampling, and a GLM-0.5B language decoder.