API Authentication
Require authentication before allowing clients to access protected AI inference endpoints.
Prevent unauthorized clients from directly interacting with the AI model.
Modern organizations are increasingly deploying Edge AI and locally hosted Large Language Models (LLMs) to perform AI inference directly within their own infrastructure.
Local AI platforms allow organizations to run AI models without sending every inference request to an external cloud-based AI service. However, exposing AI inference APIs without appropriate authentication and access controls can allow unauthorized users or applications to interact with the AI model.
In this use case, Ollama is deployed as the controlled local AI inference application on an Ubuntu Linux virtual machine inside an isolated VirtualBox laboratory.
Ollama provides a local API through which applications can interact with locally hosted AI models.
A controlled unauthorized API-access attack is simulated from Kali Linux against the Ollama inference API.
The assessment evaluates whether an unauthorized client can access the AI inference endpoint and submit inference requests without proper authorization.
OWASP ZAP is used to inspect API communication, while cURL is used to perform controlled API requests. Wireshark is used for network analysis. Wazuh is used for security monitoring, and OpenSearch is used for centralized investigation.
After identifying the API-access weakness, authentication and endpoint-access controls are implemented. The same unauthorized-access test is repeated to verify that unauthorized inference requests are rejected while legitimate AI applications continue to use the Ollama service.
Ollama is the real open-source application used in this project.
It allows locally hosted AI models to be executed and accessed through an API.
An AI inference API should not be exposed to unauthorized clients. If an Ollama API endpoint is accessible without appropriate network restrictions or authentication controls, an unauthorized user may be able to send inference requests to the locally hosted AI model.
The security problem is therefore:
The controlled attack evaluates whether an unauthorized laboratory client can access the Ollama API and submit inference requests without the required authorization. The attack is performed exclusively against the locally deployed Ollama environment. No external AI services, production systems, or real organizational data are involved.
The primary security concept is secure access control for locally hosted AI inference APIs.
AI inference endpoints should accept requests only from authorized clients. The security assessment evaluates client identity, API endpoint exposure, authentication state, request source, inference requests, request timestamps, API responses, security events, and access-control decisions.
The secure processing flow is:
Require authentication before allowing clients to access protected AI inference endpoints.
Prevent unauthorized clients from directly interacting with the AI model.
Restrict access to the Ollama inference API to authorized clients.
Prevent unauthorized applications or users from accessing the model endpoint.
Restrict access to the AI API to trusted systems or network segments.
Reduce unnecessary exposure of the local inference service.
Validate incoming inference requests before processing them.
Prevent invalid or unauthorized requests from reaching the AI model.
Avoid exposing the Ollama API beyond the systems that require access.
Reduce the attack surface of the local AI infrastructure.
Monitor inference requests and API activity.
Detect unusual or unauthorized AI API usage.
Analyze network communication involving the AI inference service.
Identify unexpected clients and abnormal API access patterns.
Record relevant API-access and authentication events.
Support detection and investigation of unauthorized AI service access.
Use Wazuh and OpenSearch to investigate monitored events.
Correlate AI API activity and establish an incident timeline.
Repeat the unauthorized API-access test after implementing security controls.
Confirm that unauthorized clients can no longer access the protected inference endpoint.
Ollama is the real open-source local AI application used in this project.
OWASP ZAP is used to inspect API communication.
cURL is used from Kali Linux to perform controlled API requests.
Wireshark is used to analyze network communication involving the Ollama API.
Wazuh is used for security monitoring.
OpenSearch is used to investigate security events collected through Wazuh.
Kali Linux is used as the authorized security-testing environment.
Ubuntu hosts the Ollama environment.
VirtualBox provides the isolated cybersecurity laboratory.