OCR engines need roughly 20 to 30 pixels of character height to classify reliably, and that single number, not megapixel count or DPI, sets camera resolution for a smart camera text-reading job [S1]. Below 20 px the engine starts dropping small or unusual fonts, and above about 40 px the sensor is paying for detail the recognition stage discards [S1]. Treat 25 pixels of x-height as the working design target, then validate against the actual engine and font in use [S1][S4].
The arithmetic is one line: sensor vertical pixels, divided by field-of-view height, multiplied by character height, gives pixels per character [S1]. A 13MP module with 3000 vertical pixels over a 297 mm A4 page height yields only 10 pixels on a 1 mm x-height glyph, while the same module over a 54 mm ID-card field-of-view yields about 55 pixels on the same glyph [S1]. That gap is why projects that buy resolution on the data sheet still fail in the field: lens and working distance spend the pixels before the sensor ever sees the character [S1].
Pixel-Height Baselines Across Western, Asian, and LPR Targets
Western (Latin, Greek, Cyrillic) OCR performs best at 25 to 30 px character height, which a 300 to 400 DPI scan of a 12 pt Times New Roman page satisfies at 78 characters per line and 46 lines per page [S4]. Asian (Japanese, Chinese, Korean) OCR shifts the target to roughly 48 by 48 pixels per character at 300 DPI for a 12 pt body font, with 30 by 30 pixels per character as the acceptable minimum, and 400 to 600 DPI recommended below 10.5 pt [S4]. License-plate recognition uses a different geometry: roughly 100 pixels across plate width is a practical design target for reliable ALPR, a wider spec because the engine reads whole plate strings rather than per-glyph stroke detail [S6].
These three baselines share one rule: pixels scale with the smallest feature, not with the document. A passport MRZ line, a 1D barcode, and a 6 pt lot-code label all live or die on the same pixels-per-feature calculation, and a 6 MP industrial camera with the wrong lens loses to a 2 MP module with the right field of view [S1]. For broader machine-vision context, the smallest relevant feature (a scratch, edge, code, dimensional difference, or OCR character) is the input that drives every downstream decision including sensor type, shutter, and interface [S5].
From Pixel Count to Megapixel and DPI: The Math Engineers Skip
Convert the pixel-height target to a sensor spec by rearranging the field-of-view equation: required vertical pixels equals (FOV height divided by character height) multiplied by 25 [S1]. A conveyor reading 8 mm tall date codes over a 400 mm wide FOV needs 400 / 8 times 25, or 1250 vertical pixels minimum, which a 1.6 MP global-shutter sensor covers with margin [S1][S5]. The 300 DPI baseline for document OCR is the same idea in scanner units: roughly 5 pixels per mm, which a 25 mm tall character hit on a 5 MP sensor at 200 mm working distance matches to within a few percent [S1].
DPI and pixels-per-character are not interchangeable, and confusing them is the second most common OCR sizing error [S1]. DPI is a scan-resolution unit tied to the physical page; pixels per character is a recognition-engine unit tied to the glyph, and the bridge between them is the x-height of the font in mm, not the document size in inches [S1][S4]. That is why an A4 scanner set to 600 DPI can still fail on 6 pt footnotes: the DPI is high, but the x-height in mm at 6 pt is about 1.5 mm, and 600 DPI gives 600 / 25.4 times 1.5, or roughly 35 pixels, which is acceptable but tight [S1][S4].
Monochrome vs Colour, Shutter, and Dynamic Range: The Three Specs That Break OCR

Monochrome sensors use every pixel for spatial detail, while single-chip Bayer colour sensors reconstruct two colour channels per pixel through interpolation, lowering effective detail resolution and reducing light sensitivity per pixel [S5]. For OCR on low-contrast, laser-etched, or dot-peen marks, monochrome is almost always the safer pick, and a colour camera typically needs more nominal pixels to match a mono camera on the same character height target [S5]. The trade-off shows up in dark-on-dark, reflective, or coloured-background text, where colour separation would help but monochrome with controlled lighting usually wins on raw stroke contrast [S5][S8].
Shutter type is the second spec that breaks OCR silently. A rolling shutter reads lines sequentially, so a part moving on a conveyor skews vertical strokes, and a 25 px glyph can land as a 25 px parallelogram that the engine reads as a different character [S5]. Global shutter exposes all pixels at once, preserving straight strokes at line speeds above roughly 0.5 m/s, and is the safer default for any moving target, pick-and-place cell, or robot-guided readout [S5]. Dynamic range is the third: glare, specular highlights, and uneven illumination clip stroke edges, and a clipped top or bottom of a 25 px x-height is enough to drop the read rate by an order of magnitude on reflective packaging [S1].
Working Distance, Lens, and the Field-of-View Trap
Working distance and lens choice spend the resolution before the sensor sees it, and a 13MP module with the wrong lens is the same camera as a 2MP module for OCR purposes [S1]. For a date-code label 8 mm tall on a 200 mm wide conveyor section, an 8 mm focal length lens at 400 mm working distance puts roughly 5 pixels per mm on the sensor, giving 40 pixels on the 8 mm character, above target with margin [S1][S5]. Push the same lens to 800 mm and the pixel density halves, dropping the same character to 20 px, where the read rate starts to depend on font, contrast, and engine more than on engineering [S1].
The lens FOV has to be sized to the text region, not the panel, the pallet, or the conveyor [S1]. A 50 mm lens that frames an entire 1 m wide conveyor puts 2 to 3 pixels per mm on a 5 mm code, which no engine reads reliably, while the same sensor with a 16 mm lens at the same distance puts 8 to 10 pixels per mm on the same code, well over the 25 px target for a 5 mm glyph [S1][S5]. The field of view is the variable that flips a 13 MP camera from unusable to over-specified for the same OCR job, and it is the first number to fix, before sensor, before interface, before lighting [S1].
Throughput, Lighting, and the Failure Modes That Hide Behind Pixel Counts

High-speed OCR systems typically take 30 to 1000 ms to process a two-line, 30-character image, so the camera frame rate, trigger latency, and encoder sync set the practical throughput ceiling before pixel count does [S7]. A 5 MP global-shutter camera at 60 fps reading two lines per trigger clears 120 lines per second per station, which lines up with the 30 to 1000 ms processing window on most production cells [S7]. Exposing the sensor longer than the part stays in the field of view is the same failure as the wrong lens: a 25 px character blurred to 25 px of motion smear costs the engine more than halving the pixel count would [S1][S5].
Contrast is the third leg of the stool. Dark text on light background, evenly lit, with a controlled specular angle, is the OCR reference case, and anything that breaks any of those three conditions costs pixels in the recognition stage even if the sensor is sized perfectly [S8]. Dot-peen, laser-etched, embossed, and curved-surface text all reduce the effective stroke-to-background contrast below the camera's measured contrast, and Adaptive Recognition's optimal-character window of 20 to 60 px tall reflects that wider operating envelope [S9]. The shortest path to a passing read rate is to validate the engine on a 25 px x-height image, then lock the lens, light, and shutter to keep that 25 px clean across the full inspection volume [S1][S8].
Selection Criteria: Who Needs What, and What to Skip
A 5 MP global-shutter monochrome camera with an 8 to 16 mm lens, 3000 to 4000 vertical pixels over the text region, and a 25 px x-height target is the right starting point for most factory OCR, date-code, and 1D/2D label jobs [S1][S5]. For passport, MRZ, and ID-card readers, the same pixel-height rule applies, but the field of view is the card, and 50 to 55 pixels per 1 mm glyph is achievable with a 6 to 13 MP sensor on a 12 to 25 mm lens at 150 to 300 mm working distance [S1]. For ALPR, target 100 px across plate width with a global shutter and pulsed IR illumination, and treat plate pitch, not plate height, as the design feature [S6].
Skip colour unless the text colour is part of the inspection, skip rolling shutter for any moving target above roughly 0.5 m/s, and skip resolutions above 8 MP unless the field of view is wider than 300 mm, because the extra pixels land on background, not on characters [S5]. Webcams, action cameras, and phone-grade sensors fail OCR not on pixel count but on aggressive noise reduction that smears the 1 to 2 pixel stroke transitions the engine depends on [S1]. The right camera is the one that puts clean, undistorted 25 px x-heights on the smallest character in the inspection, with margin for contrast, motion, and lighting variation, and that is a system spec, not a sensor pick [S1][S5].
Tracking signals for the next 6 months: OEM OCR engines publishing per-font minimum pixel-height tables rather than blanket DPI recommendations, and global-shutter 5 MP modules under USD 250 landing in mid-range smart-camera lines, both of which would shift the cost side of the OCR sizing calculation.
For component-level specifications, see height gauge, and smart meter.