Stage 1 – API Request Handling
Clients send API requests to the API gateway, which receives and manages incoming traffic.
Develop the API gateway service using FastAPI and package it as a Docker container.
This project focuses on optimizing the cloud resources used by an API Gateway and Service Management Application. The application manages API requests and routes traffic between clients and backend services. The architecture continuously monitors resource usage, identifies over-provisioned or underutilized resources, and adjusts CPU, memory, and service capacity according to actual workload requirements. The main focus is to use the required cloud resources efficiently while maintaining API performance and service availability.
To design a resource rightsizing and usage optimization architecture that reduces unnecessary cloud resource consumption while maintaining the performance and availability of an API Gateway and Service Management Application.
Clients send API requests to the API gateway, which receives and manages incoming traffic.
Develop the API gateway service using FastAPI and package it as a Docker container.
The API gateway validates requests and routes them to the required backend services.
Configure API routes and deploy the gateway and backend services as Kubernetes workloads.
CPU, memory, network, request rate, and service utilization are continuously collected.
Configure Prometheus to collect workload metrics and Grafana to display resource utilization dashboards.
Collected metrics are analyzed to identify services that are over-provisioned or consuming fewer resources than allocated.
Use Python to analyze resource usage against configured CPU and memory requests and limits.
Resource allocations are adjusted according to actual workload requirements.
Modify Kubernetes CPU and memory requests/limits based on analyzed utilization. Kubernetes supports defining resource requests and limits for containers.
When API traffic increases, additional service replicas are created; when traffic decreases, unnecessary replicas are reduced.
Configure Kubernetes Horizontal Pod Autoscaler using resource or application metrics. HPA automatically adjusts workload replicas according to observed demand.
The optimized resource configuration is continuously monitored to verify that performance is maintained and unnecessary resource usage is reduced.
Compare resource utilization before and after optimization and continuously review the configuration.
Provides compute capacity for running the application and Kubernetes workloads.
Provides isolated networking for the API gateway and backend services.
Provides persistent block storage for workloads that require local persistent data.
Stores configuration files, reports, logs, and optimization history.
Packages the API gateway and backend services into portable containers.
Deploys, manages, scales, and controls resource allocation for application workloads.
Collects CPU, memory, network, and application workload metrics.
Provides dashboards for analyzing API traffic and resource utilization.
Automates provisioning and modification of cloud infrastructure.
Automates server configuration and application environment management.
Controls access to cloud resources and services.
Controls inbound and outbound network traffic for application resources.
The proposed solution continuously monitors the API Gateway and Service Management Application using Prometheus and Grafana. Python analyzes the collected usage data to identify inefficient resource allocation. Kubernetes resource requests and limits are adjusted according to actual requirements, while Horizontal Pod Autoscaler can dynamically increase or decrease service replicas based on workload metrics. OpenTofu and Ansible automate infrastructure provisioning and configuration changes. This approach helps the application use only the resources required for its current workload while maintaining API performance and availability.