Article Views: 29
While multi-service architectures are increasing in popularity, the financial industry is being challenged as regulatory requirements continue to grow and customers’ expectations for seamless services have never been higher. Current observability practices, which largely depend on the threshold-based alerting and on reactions, are not efficient in the identification of complex failure patterns and in the prevention of cascading degradation. This article presents a new three-tier AI-driven framework to shift from reactive monitoring to predictive resilience with composite anomaly detection, a failure forecasting component, and trace-aware RCA. In a controlled synthetic evaluation, the framework demonstrated strong anomaly-detection performance (MCC: 0.995, ROC-AUC: 0.9998), effective root cause classification across 35 different error types (F1-Macro: 0.8304), and revealed significant challenges in failure prediction (MCC: 0.019). The framework demonstrated measurable anomaly-detection benefits within a controlled synthetic evaluation while providing explainable outputs that may support operational decision-making and compliance alignment.
AIOps; Predictive Resilience; Anomaly Detection; Microservices; Enterprise Architecture; Observability Governance.