BACK

Multi-Region Resilience for a Compliant AI SaaS Platform

The customer is a fast-growing AI SaaS company delivering a generative-AI platform to regulated industries, including financial services, alongside an ERP product for MSMEs with vision- and audio-based order-capture capabilities. Its platforms harness advanced generative AI to provide intelligent services while maintaining strict security, privacy, and compliance requirements, and are offered in both standard SaaS and Enterprise SaaS models with strong multi-tenant data isolation. As the customer prepared to onboard several enterprise customers, it needed to move off its existing third-party GPU-cloud infrastructure onto a production-ready, compliant AWS foundation — and then make its critical workloads resilient enough to meet stringent business-continuity and data-sovereignty obligations, without disrupting its growth roadmap. Aivar partnered with the customer to migrate to AWS and engineer a resilient, multi-region architecture anchored in India.
The customer is a fast-growing AI SaaS company delivering a generative-AI platform to regulated industries, including financial services, alongside an ERP product for MSMEs with vision- and audio-based order-capture capabilities. Its platforms harness advanced generative AI to provide intelligent services while maintaining strict security, privacy, and compliance requirements, and are offered in both standard SaaS and Enterprise SaaS models with strong multi-tenant data isolation. As the customer prepared to onboard several enterprise customers, it needed to move off its existing third-party GPU-cloud infrastructure onto a production-ready, compliant AWS foundation — and then make its critical workloads resilient enough to meet stringent business-continuity and data-sovereignty obligations, without disrupting its growth roadmap. Aivar partnered with the customer to migrate to AWS and engineer a resilient, multi-region architecture anchored in India.
No items found.

Customer Challenge

Running on third-party GPU-cloud hosting, the customer could not offer the availability, data-sovereignty controls, and disaster-recovery posture that enterprise customers in regulated sectors require. The engagement had to deliver a production-grade AWS platform and make it resilient to regional disruption. Key challenges included:

‍

  • Migration off GPU-Cloud Hosting: Transitioning containerized applications and stateful services (PostgreSQL, Redis, Kafka, Milvus) to AWS while preserving performance, security, and reliability.
    ‍
  • Regulated-Industry Resilience: Meeting SEBI Business Continuity Planning (BCP/DR) expectations — disaster declaration within 15 minutes, RTO ≤ 30 minutes, and RPO ≤ 5 minutes for critical systems.
    ‍
  • Data Sovereignty and Compliance: Keeping all PII and model processing within India regions to satisfy the DPDP Act, constraining both the primary and disaster-recovery regions to India.
    ‍
  • Highly Available Stateful Services: Deploying databases and messaging as resilient, self-healing, operator-managed workloads with automated backup and recovery.
    ‍
  • Multi-Tenant Isolation at Scale: Guaranteeing tenant data isolation while scaling to a growing enterprise customer base.
    ‍
  • Cost-Efficient Resilience: Meeting sub-hour recovery objectives without the steady-state cost of a full active-active second region.

Solution

Aivar migrated the customer from its GPU-cloud hosting to AWS and engineered a compliant, resilient platform — a production environment in Mumbai plus an Active-SemiActive high-availability tier spanning Mumbai and Hyderabad, built on AWS best practices and well-architected principles:
‍

  • Secure AWS Foundation: AWS Control Tower multi-account structure, least-privilege IAM, and Security Hub, GuardDuty, and Config with centralized logging, VPC design, and AWS WAF.
    ‍
  • Container Orchestration: Amazon EKS multi-AZ clusters for the application platforms.
    ‍
  • Resilient Stateful Services: PostgreSQL via Crunchy PGO (automated snapshots and point-in-time recovery), Kafka via Strimzi with MirrorMaker 2, Milvus vector database, and Redis via the Opstree operator.
    ‍
  • Active-SemiActive Multi-Region DR: Mumbai (ap-south-1) active and Hyderabad (ap-south-2) semi-active, scaling to full capacity only on failover — targeting RTO ≤ 30 min and RPO ≤ 5 min in line with SEBI BCP/DR.
    ‍
  • Cross-Region Replication: PostgreSQL cross-region replication, Kafka bidirectional replication (MirrorMaker 2), Milvus application-driven parallel ingestion, and Amazon S3 Cross-Region Replication.
    ‍
  • Automated Failover: Amazon Route 53 health checks and DNS failover, VPC peering between regions, and a GitHub Actions workflow that scales the Hyderabad cluster on disaster declaration.
    ‍
  • Generative AI, In-Region: Amazon Bedrock for generative-AI capabilities kept within India to preserve DPDP data residency.
    ‍
  • DevSecOps and IaC: Terraform-only Infrastructure as Code, Argo CD GitOps with drift detection, and blue-green, zero-downtime production deployments.
    ‍
  • Observability: OpenTelemetry → Prometheus + Thanos → Grafana delivering SLA metrics and failure detection.

Architecture

The target AWS architecture was designed for high availability, data sovereignty, and cost-efficient disaster recovery, comprising:
‍

  • AWS Control Tower multi-account foundation for governance and compliance
    ‍
  • Amazon EKS multi-AZ clusters in Mumbai (active) and Hyderabad (semi-active)
    ‍
  • PostgreSQL (Crunchy PGO) with automated snapshots, point-in-time recovery, and cross-region replication
    ‍
  • Kafka (Strimzi + MirrorMaker 2) for bidirectional cross-region streaming (RPO ≤ 5 min)
    ‍
  • Milvus vector database with application-driven parallel ingestion to both regions
    ‍
  • Redis (Opstree operator) for caching and session management
    ‍
  • Amazon S3 encrypted object storage with Cross-Region Replication (CRR)
    ‍
  • Amazon Route 53 health checks and DNS failover; VPC peering between regions
    ‍
  • Amazon Bedrock for in-region generative-AI capabilities
    ‍
  • Terraform IaC, Argo CD GitOps, and OpenTelemetry/Prometheus/Grafana observability
    ‍
Anonymized resilience architecture across Mumbai and Hyderabad
Active-SemiActive multi-region resilience architecture (Mumbai active, Hyderabad semi-active).

Key Outcomes

  • Regulatory-Grade Resilience: Active-SemiActive multi-region design targeting RTO ≤ 30 minutes and RPO ≤ 5 minutes, aligned with SEBI BCP/DR guidelines.
    ‍
  • Data Sovereignty Assured: Both primary and disaster-recovery regions anchored in India, meeting DPDP Act requirements.
    ‍
  • Production-Ready Platform: Fully migrated off third-party GPU-cloud hosting to a scalable, multi-tenant AWS environment ready for enterprise onboarding.
    ‍
  • Cost-Efficient Disaster Recovery: Semi-active secondary region scales to full capacity only on failover, controlling standby compute cost.
    ‍
  • Automated, Zero-Downtime Operations: Terraform IaC, Argo CD GitOps, and blue-green deployments across all production components.
    ‍
  • Operational Independence: Documented, tested failover and fallback runbooks with full observability, enabling the customer to operate the solution independently.

Explore Other Case Studies