Anomaly detection answers a binary question (normal or not) on imbalanced data, while defect classification assigns a labeled type to every observation on balanced classes; choosing the wrong one is the single most common reason visual inspection pilots fail to move past the lab [S2][S3].
The 2025 industrial anomaly detection market reached $6.90B and is projected to grow to $28B by 2034 at a 16.83% CAGR, with the machine-learning and AI subsegment tracking a faster 18.92% CAGR, and ML/AI already commands roughly 59% of that spend [S3]. In parallel, the May 2025 arXiv paper "Detect, Classify, Act" introduced the MVTec-AC and VisA-AC benchmarks, reporting 80.4% accuracy on MVTec-AD and 84% on MVTec-AC with a two-stage LLM-based pipeline, beating prior baselines by 5 percentage points [S4].
Defining the two tasks and the gap between them
Anomaly detection, as framed by Steinwart et al. (JMLR 2005), treats anomalies as points outside the dense level sets of the data-generating distribution, formalized through density thresholds ρ on a reference measure μ; this binary framing has underpinned the field for two decades and is the basis of the standard one-class SVM [S5].
Defect classification, by contrast, is a supervised multiclass problem that requires a balanced labeled set covering every defect category the system is expected to recognize; an open dataset that the model has never seen is by definition outside its scope [S2][S3]. The May 2025 arXiv pipeline highlights exactly this gap: visual anomaly detectors such as PatchCore or DDAD can segment a scratch at sub-100ms latency but cannot tell you whether that scratch is a fatal crack, a cosmetic surface mark, or a benign color change from a design revision [S4].
Algorithm categories by data regime
Supervised anomaly detection collapses to plain classification: both normal and anomalous samples are labeled, the dataset is imbalanced, and the model is a binary classifier; in that regime a 97% accuracy on a 97/3 split is no better than predicting the majority class, a common pitfall illustrated in the start-up survival example [S2].
Semi-supervised anomaly detection trains on a clean normal set and scores new points by reconstruction error (autoencoders) or distance from the learned manifold (one-class SVM, PatchCore); this is the dominant regime in modern vision inspection because labeling every defect mode is impractical [S5][S7]. Unsupervised methods such as DBSCAN clustering of autoencoder residuals or Isolation Forest assume anomalies are rare and unbounded in type, which works on noisy shop-floor data but produces no defect taxonomy [S1][S3].
Criteria-based comparison: anomaly detection vs defect classification

On label requirement, anomaly detection tolerates zero or partial labels (autoencoder, one-class SVM, PatchCore) while defect classification demands a fully labeled balanced set per category; the Klarák (PMC 2024) U2S-CNN pipeline on MVTec gear wheels reported 108 correctly detected regions out of 177 candidates, with 69 mislabels, illustrating the cost of mixing the two regimes without enough labeled defect examples [S1].
On inference cost, single-step detectors such as YOLO and SSD deliver the 30–60 FPS minimum that production lines require but at the price of lower localization precision, whereas two-step detectors (R-CNN, Faster R-CNN, Mask R-CNN, Cascade R-CNN) trade compute for tighter bounding boxes and pixel-level masks [S1]. On output semantics, anomaly detectors return a score and a heatmap; defect classifiers return a class ID and a confidence, which is the format MES, SCADA, and traceability systems actually consume; see related vision-system sensor selection logic for how this maps to upstream measurement hardware.
Where each method fits on the plant floor
Anomaly detection is the right tool for new product introduction, low-volume SKUs, and any line where defect modes evolve faster than the labeling team can keep up; the U2S-CNN autoencoder-to-DBSCAN-to-Xception chain on MVTec gear wheels, with 177 detected regions against 205 actual damaged areas, shows the architecture is mature enough to localize defects in unsupervised mode before any classifier is bolted on [S1].
Defect classification is the right tool for high-volume mature lines where the defect catalog is stable and the cost of a misroute (a cosmetic mark sent to scrap, a real crack shipped to the customer) is quantifiable; MVTec-AC and VisA-AC were published precisely to give practitioners a labeled benchmark for this regime, and the VELM pipeline's 80.4% / 84% numbers set the 2025 state of the art on those benchmarks [S4]. For tactile or dimensional inspection upstream of a vision cell, see displacement sensor selection on wear and drift to align sensor bandwidth with classifier input rate.
Failure modes and limits to budget for

Anomaly detection fails on benign novelty: design-revision color changes, scheduled tooling marks, and lighting drift all trigger false rejects, which is why the VELM pipeline routes every flagged region through a multimodal LLM to decide whether it is admissible [S4]. Conversely, a defect classifier fails silently on out-of-catalog defects; it returns the closest trained class with high confidence, which is worse than a flag because the downstream system trusts the label.
Hybrid architectures mitigate both: Klarák's U2S-CNN layers unsupervised localization (autoencoder), unsupervised grouping (DBSCAN), and supervised labeling (Xception) and reports 108/177 correct region-level decisions, a recall profile the authors frame as a proof of concept rather than a production-ready figure [S1]. MVTec's reference pipeline similarly documents that deep-learning anomaly detection is the only path to automated surface inspection with pixel-precise localization, but only when paired with a downstream decision layer that knows defect semantics [S7].
Standards, sourcing, and procurement signals
There is no single IEC or ISO standard that prescribes anomaly detection vs classification for industrial vision; the relevant guidance is the MVTec-AD and MVTec-AC benchmarks and the 2025 arXiv evaluation protocol from the University of Freiburg and Endress+Hauser, which together define what "good" looks like for a published model [S4]. Procurement signals worth tracking through 2026 are (a) VELM-style two-stage pipelines (unsupervised detector plus LLM-based classifier) moving from arXiv to vendor datasheets, and (b) MVTec-AC and VisA-AC label sets being adopted as de facto reference catalogs the same way MVTec-AD has been since 2019.
For plants that already run a vision cell tied to a flow-meter or PLC-integrated quality gate, the practical trackable signal is whether your vendor publishes per-class F1 on a named benchmark; if they only publish AUROC on a private set, you are buying anomaly detection, not defect classification, and your MES rules must be written accordingly.
For component-level specifications, see gas detection, and pressure transmitter.