BforBank: Performance & 360° Observability

CASE STUDY · ONLINE BANKING

Securing the scale-up to 200,000 active users.

Adservio structures a cross-functional Performance & Observability team for the Crédit Agricole group's online bank: monitoring, peak anticipation, flawless client experience.

BforBank : Adservio case study
Client
BforBank (Crédit Agricole group)
Expertise
Augmented SRE · 360° Observability · Performance Engineering
Tech Stack
Datadog · Grafana · OpenTelemetry · K6 · AWS · Kubernetes
Engagement
8 experts · 24 months
CONTEXT

Project Context

BforBank aims to reach 200,000 active users across its online banking journeys. This scale-up requires a complete overhaul of observability and performance practices: the legacy tooling, inherited from the group CIO function, can no longer correlate signals across infrastructure, applications and client experience.

Without a cross-functional framework, the DevOps and application teams struggled to correlate signals across technical layers. Each incident triggered a multi-team crisis meeting, each user degradation was detected several minutes late: a major operational and reputational risk in a banking market where trust is measured in milliseconds.

The challenge: moving from a reactive approach to proactive monitoring able to detect drifts before they impact the end user, and embedding a lasting SRE culture across the 12 product teams, aligned with Google SRE standards and adapted to banking regulatory constraints (DORA, ACPR).

Strategic Objectives

(01)

360° Observability

Implement unified monitoring covering infrastructure, applications, network and end-user experience: no more per-layer silos, replaced by a correlated view per business journey (login, transfer, card payment).

(02)

Performance Engineering

Industrialise load testing within CI/CD, with automated non-regression thresholds on every release. Anticipate scaling to 200,000 active users without degrading the SLOs of critical journeys.

(03)

SRE culture

Spread SRE practices across the 12 product teams: measurable SLI/SLOs, error budgets negotiated with the business, blameless postmortems, regular game days, moving from a reactive stance to an operational discipline.

Solutions Delivered by Adservio

Adservio deployed a multidisciplinary SRE squad (SRE Lead, Performance Engineers, Data Reliability, FinOps) to structure the practice over 24 months.

(01)

SLO mapping

Definition of critical indicators per user journey (login, transfer, payment, KYC, mobile) with error budgets negotiated directly with product management. 48 active SLOs aligned with real usage.

(02)

Unified observability platform

Rollout of a shared Datadog / Grafana / OpenTelemetry stack at group scale, with per-journey dashboards, intelligent alerting (multi-window burn rate) and automatic infra/app/RUM correlation.

(03)

Industrialised load testing

Automation of performance tests (K6, JMeter, Gatling) within CI/CD with automated reporting, blocking non-regression thresholds and quarterly capacity planning: performance becomes a quality gate, not an annual audit.

(04)

Driving the SRE culture

Training the 12 product teams on SLI/SLOs and error budgets, running monthly game days and blameless postmortems, evangelising within the group CIO function with the creation of an internal SRE Academy.

(05)

FinOps & green IT

Continuous optimisation of AWS cloud consumption (rightsizing, autoscaling, spot) coupled with carbon tracking per business journey: a 22% reduction in the infrastructure bill while absorbing user growth.

Measurable SLOs. A managed error budget.
A drift detected in under 5 minutes.

Critical journeys are defined in an auditable YAML file. The Datadog + OpenTelemetry engine executes, computes the burn rate and triggers a deployment freeze when the budget is exhausted: 200,000 users monitored in real time.

Results

×4
User capacity

Scaling up to 200,000 active users with no product freeze.

99.99%
SLO achieved

Availability sustained on payment, login and transfer journeys.

÷3
MTTR P1

Mean time to remediation on P1 incidents cut from 28 to 9 minutes.

÷5
Time-to-detect

Business incident detection cut from 25 minutes to under 5.

48
Active SLOs

Critical indicators aligned with real usage per journey.

−22%
Cloud costs

AWS bill reduced despite the ×4 growth in usage.

Impact

User capacity ×4

Platforms able to support scaling up to 200,000 active users, with SLOs maintained across all critical journeys. No more product freezes caused by performance during marketing campaigns.

99.99% SLO achieved

Availability targets met on the payment, login and transfer journeys. Client commitments hold under all peaks, including during sales periods and month-ends.

MTTR cut threefold

Faster incident detection and resolution thanks to unified observability and proactive alerts. Mean time to remediation brought down from 28 to 9 minutes on P1 incidents.

SRE culture embedded

Cross-functional adoption by 12 product teams, training on SLI/SLOs and error budgets, and the rollout of an internal SRE Academy. Operational discipline is now embedded in daily work.

FinOps & green IT

22% lower AWS cloud costs despite a ×4 growth in usage, and a measured 18% drop in carbon footprint per user journey: reliability no longer comes at the cost of growth.

Time-to-detect cut fivefold

Before the programme, a business incident could take up to 25 minutes to be detected. Today, alerting based on SLOs and burn rate identifies it in under 5 minutes, often before the end user.

TALK TO AN EXPERT

Secure your scale-up with Adservio

Let's discuss your SLO, observability and performance challenges. An Adservio expert gets back to you within 24h.

By submitting this form, you agree to our privacy policy.