Master's Thesis from the year 2025 in the subject Mathematics - Applied Mathematics, grade: 1,0, Justus-Liebig-University Giessen, language: English, abstract: This thesis investigates whether machine learning (ML) - i.e., statistical learning - can provide a scalable alternative for detecting errors in official statistics. Following a mathematical framework, we formalize ML classifiers from four algorithmic families for their deployment in the detection task: Random Forests, Boosting Methods, Support Vector Machines, and Neural Networks - and train and evaluate them on data provided by the German Federal Statistical Office. To address the scarcity of labeled data, we implement a systematic noise injection strategy that exposes models to controlled error rates and rare violation types. Furthermore, we extend the investigation beyond performance evaluation to examine the robustness of the models when the underlying stochastic mechanisms behind faulty data change. The results demonstrate that especially tree-based ensemble methods can achieve ROC- and PR-AUC scores exceeding 0.98, nearly matching deterministic validation and scaling linearly with computational efficiency of O(n). However, an important finding emerges: while classifiers remain robust when error probabilities shift from uniform randomness to covariate-dependent patterns, performance degrades when errors depend on the variable's own values. In particular, neural networks experience a decline in PR-AUC from 0.9 to 0.7. This work addresses a gap in the literature by providing the first empirical evidence of the impact of stochastic error-generation mechanisms on ML validation performance. The findings suggest that although the deployment of artificial intelligence in data validation offers an efficient alternative to deterministic validation, there are considerations related to the stochastic mechanism behind the data that must be acknowledged, as they can affect model performance.
Die Inhaltsangabe kann sich auf eine andere Ausgabe dieses Titels beziehen.
Anbieter: California Books, Miami, FL, USA
Zustand: New. Bestandsnummer des Verkäufers I-9783389198377
Anzahl: Mehr als 20 verfügbar
Anbieter: BuchWeltWeit Ludwig Meier e.K., Bergisch Gladbach, Deutschland
Taschenbuch. Zustand: Neu. This item is printed on demand - it takes 3-4 days longer - Neuware 80 pp. Englisch. Bestandsnummer des Verkäufers 9783389198377
Anzahl: 2 verfügbar
Anbieter: buchversandmimpf2000, Emtmannsberg, BAYE, Deutschland
Taschenbuch. Zustand: Neu. This item is printed on demand - Print on Demand Titel. Neuware 80 pp. Englisch. Bestandsnummer des Verkäufers 9783389198377
Anzahl: 1 verfügbar
Anbieter: AHA-BUCH GmbH, Einbeck, Deutschland
Taschenbuch. Zustand: Neu. Druck auf Anfrage Neuware - Printed after ordering. Bestandsnummer des Verkäufers 9783389198377
Anzahl: 1 verfügbar
Anbieter: preigu, Osnabrück, Deutschland
Taschenbuch. Zustand: Neu. Enhancing Statistical Editing in Official Statistics | Validation and Detection of Faulty Data through Artificial Intelligence | Carlos Andres Salamanca Dávila | Taschenbuch | Englisch | 2026 | GRIN Verlag | EAN 9783389198377 | Verantwortliche Person für die EU: preigu GmbH & Co. KG, Lengericher Landstr. 19, 49078 Osnabrück, mail[at]preigu[dot]de | Anbieter: preigu Print on Demand. Bestandsnummer des Verkäufers 135930035
Anzahl: 5 verfügbar