GenAI Deployment on OCI: Strategies for Scalable and Production-Ready AI Applications
You have built a brilliant GenAI model. It answers questions perfectly. It understands context. It is amazing — on your laptop. 💻
But here is the hard truth that nobody warns you about: Building the AI is only half the job. Getting it running reliably for thousands of users, every day, without breaking — that is the other half.
In the real world, 70% of AI projects fail NOT because the model was bad — but because nobody knew how to deploy, scale, and manage it properly. Deployment Strategy is the skill that separates a hobby project from a real product.
🏢 Our Real-World App: The SmartHR Assistant
Let us design a real app so everything we learn has a purpose. Meet SmartHR!
SmartHR is an AI-powered chatbot for a company with 10,000 employees. Employees can ask it questions like:
- 💬 "How many vacation days do I get per year?"
- 💬 "What is the maternity leave policy?"
- 💬 "How do I apply for a salary advance?"
Instead of waiting for an HR email reply (which can take days!), SmartHR answers in under 2 seconds — 24 hours a day, 7 days a week. 🌟
THE SMARTHR APP — HIGH-LEVEL DESIGN ┌─────────────────────────────────────────────────────────────────────────┐ │ │ │ 👩💼 Employee │ │ (Web / Mobile App) │ │ │ │ │ │ "How many vacation days do I get?" │ │ ▼ │ │ ┌─────────────────────────────┐ │ │ │ OCI API Gateway │ ◄── Single entry point for all users │ │ │ (The front door 🚪) │ Routes requests, checks security │ │ └──────────────┬──────────────┘ │ │ │ │ │ ┌──────────┴──────────┐ │ │ │ │ │ │ ▼ ▼ │ │ ┌──────────┐ ┌──────────────────┐ │ │ │ OCI │ │ OCI Container │ │ │ │ Functions│ │ Instances │ │ │ │ (Quick │ │ (Heavy workloads │ │ │ │ replies)│ │ like Vision AI) │ │ │ └────┬─────┘ └────────┬─────────┘ │ │ │ │ │ │ └───────────┬─────────────┘ │ │ ▼ │ │ ┌─────────────────┐ │ │ │ OCI GenAI │ │ │ │ (Cohere / Llama │ │ │ │ answers the │ │ │ │ question) │ │ │ └─────────────────┘ │ │ │ └─────────────────────────────────────────────────────────────────────────┘
This one app will teach us every deployment concept you need to know in 2026. Let us go step by step! 👟
🚪 Deployment Option 1 — OCI Functions (Serverless)
What is "Serverless"?
Imagine you open a lemonade stand. 🍋 A traditional server is like hiring a full-time employee who stands at the counter all day — even when no customers come. You pay them 24/7 whether they are busy or sitting idle.
Serverless (OCI Functions) is like a magic employee who appears only when a customer arrives, serves them instantly, and disappears again — and you only pay for the exact seconds they worked! No customer = zero cost. ✨
OCI Functions is perfect for the SmartHR chatbot because:
- Employees ask questions randomly throughout the day (not all at once)
- Each question takes only 1–3 seconds to process
- You do not want to pay for an idle server at midnight when nobody is using it
- You need it to handle sudden spikes (like 9 AM Monday when everyone checks policies)
HOW OCI FUNCTIONS WORKS:
NORMAL TIMES (few users):
┌──────────┐ ┌────────────────────────────────────────┐
│ 1 User │────►│ 1 Function Instance wakes up, answers │
│ asks │ │ in 1.5 seconds, then sleeps again 😴 │
└──────────┘ └────────────────────────────────────────┘
💰 Cost: You pay for ~1.5 seconds of compute
BUSY TIMES (1000 users at once):
┌──────────┐ ┌─────────────────────────────────────────────────────┐
│ 1000 │────►│ OCI automatically spins up 1000 Function instances │
│ Users │ │ All answer in parallel. Done in 2 seconds. │
│ ask │ │ Then ALL 1000 instances go to sleep. 😴 │
└──────────┘ └─────────────────────────────────────────────────────┘
💰 Cost: You pay for 1000 × 2 seconds of compute
🎯 No pre-planning needed — OCI scales automatically!
📦 Deploying SmartHR on OCI Functions — Step by Step
There are 4 things you need to deploy an OCI Function:
- 📝 func.py — Your Python code (what the function does)
- 📋 func.yaml — The function's ID card (name, memory, timeout)
- 📦 requirements.txt — List of Python packages needed
- 🐳 Dockerfile — The container recipe (OCI Functions runs in Docker)
This is the brain of the SmartHR Function. When an employee asks a question, OCI runs this code. It receives the question, sends it to OCI GenAI, and returns the answer — all in one tidy Python function. 🧠
# ── FILE: func.py ──────────────────────────────────────────────
# PURPOSE: This is the OCI Function that powers SmartHR.
# It receives an employee's question, calls OCI GenAI,
# and sends back a helpful HR policy answer.
# ───────────────────────────────────────────────────────────────
import io
import json
import oci
from fdk import response
from oci.generative_ai_inference import GenerativeAiInferenceClient
from oci.generative_ai_inference.models import (
ChatDetails, OnDemandServingMode, CohereChatRequest
)
# OCI GenAI configuration
COMPARTMENT_ID = "ocid1.compartment.oc1..xxxxx"
MODEL_ID = "cohere.command-r-plus"
# The HR knowledge (in real app, this comes from Oracle 23ai Vector DB)
HR_CONTEXT = """
- Annual leave: 24 days per year for all permanent employees.
- Maternity leave: 26 weeks fully paid for female employees.
- Salary advance: Available once per year, max 2 months salary.
- Work from home: Up to 3 days per week with manager approval.
"""
def handler(ctx, data: io.BytesIO = None):
"""
OCI Functions entry point.
Every time an employee sends a question, OCI calls this function.
"""
try:
# Step 1: Read the employee's question from the request body
body = json.loads(data.getvalue())
question = body.get("question", "")
if not question:
return response.Response(
ctx, response_data=json.dumps({"error": "No question provided"}),
headers={"Content-Type": "application/json"}
)
# Step 2: Connect to OCI using Instance Principal
# (No passwords needed! OCI Functions has built-in identity)
signer = oci.auth.signers.get_resource_principals_signer()
client = GenerativeAiInferenceClient(config={}, signer=signer)
# Step 3: Build the prompt for the LLM
prompt = f"""
You are SmartHR, a friendly HR assistant.
Use only the HR Policy below to answer the question.
If the answer is not in the policy, say "Please contact HR directly."
HR POLICY:
{HR_CONTEXT}
EMPLOYEE QUESTION: {question}
ANSWER:
"""
# Step 4: Call OCI GenAI and get the answer
chat_request = CohereChatRequest(
message = prompt,
max_tokens = 300,
temperature = 0.1 # Low = factual, not creative
)
resp = client.chat(ChatDetails(
serving_mode = OnDemandServingMode(model_id=MODEL_ID),
compartment_id = COMPARTMENT_ID,
chat_request = chat_request
))
answer = resp.data.chat_response.text
# Step 5: Return the answer as JSON
result = {"question": question, "answer": answer, "status": "success"}
return response.Response(
ctx, response_data=json.dumps(result),
headers={"Content-Type": "application/json"}
)
except Exception as e:
return response.Response(
ctx, response_data=json.dumps({"error": str(e), "status": "error"}),
headers={"Content-Type": "application/json"}
)
This is the ID card for your OCI Function. It tells OCI the function's name, how much memory it needs, and how long it is allowed to run before timing out. Think of it like a recipe card for the kitchen! 🃏
# ── FILE: func.yaml ──────────────────────────────────────────── # PURPOSE: The configuration card for the SmartHR OCI Function. # ─────────────────────────────────────────────────────────────── schema_version: 20180708 name: smarthr-assistant # The function's name in OCI version: "1.0.0" # Version number (important for CI/CD!) runtime: python3.11 # We are using Python 3.11 build_image: fnproject/python:3.11-dev run_image: fnproject/python:3.11 entrypoint: /python/bin/fdk /function/func.py handler memory: 512 # 512 MB of RAM for this function timeout: 30 # Max 30 seconds to run, then timeout
These are the 4 terminal commands you type to actually deploy the function to OCI. Think of it like pressing the "Publish" button — but for cloud functions! 🚀
# ── TERMINAL COMMANDS: Deploy SmartHR to OCI Functions ─────────
# PURPOSE: Build and deploy your function to Oracle Cloud.
# Run these 4 commands in your terminal, in order.
# ───────────────────────────────────────────────────────────────
# Step 1: Tell fn CLI which OCI context to use
fn use context us-ashburn-1
# Step 2: Create an Application (a folder that holds your functions)
fn create app smarthr-app --annotation oracle.com/oci/subnetId=ocid1.subnet.oc1..xxxxx
# Step 3: Deploy the function (builds the Docker image + uploads to OCI)
fn -v deploy --app smarthr-app
# Step 4: Test it locally before going live!
echo '{"question": "How many vacation days do I get?"}' | fn invoke smarthr-app smarthr-assistant
# ✅ Expected Output:
# {"question": "How many vacation days do I get?",
# "answer": "All permanent employees receive 24 days of annual leave per year.",
# "status": "success"}
🌐 Step 2 — OCI API Gateway: The Smart Front Door
Your OCI Function is deployed and working. But right now, employees can only call it using complicated OCI URLs with security tokens. That is not user-friendly at all!
OCI API Gateway gives your SmartHR app a clean, secure, public URL that any web or mobile app can call easily. 🚪
Think of OCI API Gateway as the reception desk of a big office building. 🏢 Visitors (users) don't walk directly into every room. They talk to the receptionist first. The receptionist checks their ID (authentication), tells them which floor to go to (routing), and tracks who visited (logging). Only then do they get through. The API Gateway does all of this!
HOW OCI API GATEWAY WORKS FOR SMARTHR:
Employee's App
│
│ POST https://api.smarthr.company.com/v1/ask
│ Body: {"question": "What is maternity leave?"}
│
▼
┌─────────────────────────────────────────────────────────┐
│ OCI API GATEWAY │
│ │
│ 1. 🔒 Check: Is this user authenticated? (JWT Token) │
│ 2. 🚦 Check: Have they sent too many requests? (Rate │
│ Limiting — max 60 requests/minute) │
│ 3. 🔀 Route: Which department (tenant) are they from? │
│ HR Dept → Route to HR function │
│ Finance Dept → Route to Finance function │
│ 4. 📊 Log: Record this request for monitoring │
│ │
└──────────────────────────┬──────────────────────────────┘
│
┌─────────────┴──────────────┐
│ │
▼ ▼
OCI Function OCI Function
(SmartHR v1.0) (SmartHR v2.0 — canary)
90% of traffic 10% of traffic
(stable version) (new version testing)
⚙️ Setting Up the API Gateway — What It Looks Like
This sets up the OCI API Gateway to accept employee questions at a clean URL. It checks that the request is a POST, verifies the user's JWT token, and then forwards the request to our SmartHR OCI Function. 🚏
# ── FILE: api-deployment.yaml ──────────────────────────────────
# PURPOSE: Define the API Gateway rules for SmartHR.
# Tells OCI which URLs to accept, who is allowed in,
# and where to forward each request.
# ───────────────────────────────────────────────────────────────
displayName: SmartHR API Deployment
pathPrefix: /v1
routes:
# ── ROUTE 1: Ask HR a question ─────────────────────────────
- path: /ask
methods: [POST] # Only POST requests allowed here
# Authentication: Check the JWT token
# If no valid token → reject with 401 Unauthorized
requestPolicies:
authentication:
type: JWT_AUTHENTICATION
tokenHeader: Authorization # Token must be in the Authorization header
publicKeys:
type: REMOTE_JWKS
uri: "https://identity.oraclecloud.com/.well-known/jwks.json"
# Rate limiting: Max 60 requests per minute per user
# Protects the AI backend from being overwhelmed
rateLimiting:
rateInRequestsPerSecond: 1 # 1 per second = 60 per minute
rateKey: CLIENT_IP # Limit per individual IP address
backend:
type: ORACLE_FUNCTIONS_BACKEND
functionId: "ocid1.fnfunc.oc1..xxxxx" # Our SmartHR OCI Function
# ── ROUTE 2: Health check (no auth needed) ─────────────────
- path: /health
methods: [GET]
backend:
type: STOCK_RESPONSE_BACKEND
status: 200
body: '{"status": "SmartHR API is healthy ✅"}'
Always add a /health route to your API Gateway. Monitoring tools ping this endpoint every 30 seconds to check if your service is alive. No authentication needed — it just says "I'm alive!" 💓
🐳 Deployment Option 3 — OCI Container Instances
When is OCI Functions NOT Enough?
OCI Functions is great for fast, short tasks. But what if SmartHR needs to process an uploaded 10-page PDF — extracting text, running vision AI on embedded charts, chunking, and embedding everything?
That could take 60–120 seconds. OCI Functions has a 30-second maximum timeout. For heavy workloads, we need something more powerful — OCI Container Instances.
- ⚡ OCI Functions — Like a food delivery rider. Quick, small tasks. Pay per delivery.
- 🚛 OCI Container Instances — Like a delivery truck. Larger loads, longer journeys. Always running, but can carry much more.
WHEN TO USE WHICH: ┌──────────────────────┬───────────────────────────┬──────────────────────────┐ │ Feature │ OCI Functions │ OCI Container Instances │ ├──────────────────────┼───────────────────────────┼──────────────────────────┤ │ Startup Time │ Fast (cold start ~2s) │ Always running (~0s) │ │ Max Run Duration │ 30 seconds │ Unlimited │ │ Memory │ Up to 2 GB │ Up to 32 GB │ │ GPU Support │ ❌ No │ ✅ Yes (for heavy AI) │ │ Idle Cost │ 💰 Zero (pay per call) │ 💰 Hourly rate │ │ Use Case │ Quick Q&A (SmartHR chat) │ PDF ingestion, training │ └──────────────────────┴───────────────────────────┴──────────────────────────┘ SmartHR uses BOTH: Chat questions ──► OCI Functions (fast, cheap, auto-scales) PDF uploads ──► OCI Container (powerful, handles long tasks)
A Dockerfile is like a recipe for baking the perfect cake. It tells OCI: "Start with Python, install these packages, copy my code, and when someone runs this container — start the web server." OCI follows the recipe and creates a ready-to-run container image! 🍰
# ── FILE: Dockerfile ─────────────────────────────────────────── # PURPOSE: The recipe to build the SmartHR PDF Processor container. # This container handles large PDF ingestion jobs. # ─────────────────────────────────────────────────────────────── # Layer 1: Start with a Python 3.11 base image FROM python:3.11-slim # Layer 2: Set the working directory inside the container WORKDIR /app # Layer 3: Copy the requirements file first (for caching efficiency) COPY requirements.txt . # Layer 4: Install Python packages # --no-cache-dir keeps the image small RUN pip install --no-cache-dir -r requirements.txt # Layer 5: Copy all application code into the container COPY . . # Layer 6: Expose port 8080 so OCI can talk to this container EXPOSE 8080 # Layer 7: The command that runs when the container starts # Starts a FastAPI web server with 4 workers CMD ["uvicorn", "pdf_processor:app", "--host", "0.0.0.0", "--port", "8080", "--workers", "4"]
This is the actual PDF processing service inside the container. It exposes a REST API endpoint that receives a PDF file, extracts the text, and starts the ingestion process — exactly what happens when an HR manager uploads a new policy document! 📄➡️🧠
# ── FILE: pdf_processor.py ─────────────────────────────────────
# PURPOSE: FastAPI web server that runs inside the OCI Container.
# Receives PDF files, extracts text, creates embeddings,
# and stores them in Oracle 23ai Vector DB.
# ───────────────────────────────────────────────────────────────
from fastapi import FastAPI, UploadFile, File, BackgroundTasks
from PyPDF2 import PdfReader
from langchain.text_splitter import RecursiveCharacterTextSplitter
import io, uuid, logging
app = FastAPI(title="SmartHR PDF Processor", version="1.0.0")
logging.basicConfig(level=logging.INFO)
@app.get("/health")
async def health():
"""Health check — lets OCI know this container is alive."""
return {"status": "healthy", "service": "SmartHR PDF Processor"}
@app.post("/ingest-pdf")
async def ingest_pdf(
background_tasks: BackgroundTasks,
file: UploadFile = File(...)
):
"""
Receive an HR policy PDF, extract text, chunk it,
create embeddings, and store in Oracle 23ai.
Runs as a background task so the response is instant.
"""
# Step 1: Read the uploaded PDF bytes
pdf_bytes = await file.read()
job_id = str(uuid.uuid4())
logging.info(f"📄 Received PDF: {file.filename} | Job: {job_id}")
# Step 2: Start processing in the background
# (The user gets an instant response with job_id)
# (The actual work happens asynchronously)
background_tasks.add_task(process_pdf, pdf_bytes, file.filename, job_id)
return {
"job_id" : job_id,
"status" : "processing",
"message" : f"PDF '{file.filename}' is being processed. Check job status with job_id."
}
async def process_pdf(pdf_bytes: bytes, filename: str, job_id: str):
"""The actual PDF processing logic (runs in background)."""
# Step 1: Extract text from PDF
reader = PdfReader(io.BytesIO(pdf_bytes))
raw_text = "".join(page.extract_text() or "" for page in reader.pages)
logging.info(f"✅ Extracted {len(raw_text)} characters from {filename}")
# Step 2: Split text into chunks
splitter = RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50)
chunks = splitter.split_text(raw_text)
logging.info(f"✂️ Created {len(chunks)} chunks")
# Step 3: Create embeddings and store in Oracle 23ai
# (Calls the embedding + storage functions from previous blog post)
for i, chunk in enumerate(chunks):
# embed(chunk) → store in Oracle 23ai Vector DB
logging.info(f" Processing chunk {i+1}/{len(chunks)}")
# ... embedding and storage code here (see previous blog post)
logging.info(f"🎉 Job {job_id} complete! {len(chunks)} chunks stored.")
These commands build your container into an image, push it to OCI's container registry (like uploading to a container app store), and then launch it as an OCI Container Instance. After these commands — your PDF processor is live on the cloud! ☁️
# ── TERMINAL: Build, Push, and Deploy the Container ────────────
# PURPOSE: Get the SmartHR PDF Processor running on OCI.
# ───────────────────────────────────────────────────────────────
# Step 1: Log in to OCI Container Registry
docker login iad.ocir.io -u "tenancynamespace/your.email@company.com"
# Step 2: Build the Docker image
docker build -t iad.ocir.io/tenancynamespace/smarthr/pdf-processor:1.0.0 .
# Step 3: Push to OCI Container Registry (cloud app store)
docker push iad.ocir.io/tenancynamespace/smarthr/pdf-processor:1.0.0
# Step 4: Deploy as OCI Container Instance (using OCI CLI)
oci container-instances container-instance create \
--compartment-id "ocid1.compartment.oc1..xxxxx" \
--display-name "smarthr-pdf-processor" \
--shape "CI.Standard.E4.Flex" \
--shape-config '{"ocpus": 2, "memoryInGBs": 8}' \
--containers '[{
"imageUrl": "iad.ocir.io/tenancynamespace/smarthr/pdf-processor:1.0.0",
"displayName": "pdf-processor",
"resourceConfig": {"memoryLimitInGBs": 8, "vcpusLimit": 2}
}]'
echo "✅ SmartHR PDF Processor is live on OCI Container Instances!"
📈 Scaling AI Services for High Concurrency
Imagine it is Monday morning at 9:00 AM. 🌅 The company just released new HR policies and sent an email to all 10,000 employees. Every single person opens SmartHR at the same time to check their leave balance.
Suddenly, 10,000 requests arrive in 30 seconds. Your single server would crash. 💥 But with the right scaling strategy, SmartHR handles it effortlessly. Here is how:
SCALING STRATEGIES FOR SMARTHR (2026): ┌─────────────────────────────────────────────────────────────────────┐ │ │ │ LAYER 1: OCI Functions Auto-scaling (Chat endpoint) │ │ ───────────────────────────────────────────────────── │ │ 1 user/min ──► 1 concurrent function instance │ │ 100 users/min ──► 100 concurrent function instances (auto!) │ │ 10,000/min ──► 10,000 concurrent instances (still auto!) 🚀 │ │ OCI handles all of this — YOU don't do anything! 🙌 │ │ │ │ LAYER 2: OCI GenAI Rate Limits (Automatic management) │ │ ───────────────────────────────────────────────────── │ │ OCI GenAI Service has built-in request queuing. │ │ Burst of 10,000 requests? They queue up and process in order. │ │ No requests are dropped. ✅ │ │ │ │ LAYER 3: OCI Container Instances Autoscaling (PDF processor) │ │ ───────────────────────────────────────────────────── │ │ Low traffic ──► 1 container instance (min) │ │ High traffic ──► Up to 10 container instances (max) │ │ Scale trigger: CPU > 70% OR memory > 80% │ │ │ └─────────────────────────────────────────────────────────────────────┘
⚙️ Setting Up Autoscaling for OCI Container Instances
This tells OCI: "Keep at least 1 PDF processor running at all times. If CPU goes above 70%, add more containers automatically — up to 10 maximum. When traffic drops, shrink back down to save money." 📈📉
# ── FILE: autoscaling-policy.yaml ──────────────────────────────
# PURPOSE: Tell OCI when to add more PDF processor containers
# and when to remove them. Like a thermostat for servers! 🌡️
# ───────────────────────────────────────────────────────────────
displayName: SmartHR PDF Processor Autoscaling
# How often OCI checks if scaling is needed
coolDownInSeconds: 300 # Check every 5 minutes
# The container group to scale
resourceId: "ocid1.instancepool.oc1..xxxxx"
policies:
- policyType: threshold # Scale based on a metric threshold
# Scale UP rules:
# If CPU is above 70% → add 2 more container instances
rules:
- displayName: scale-up-on-high-cpu
action:
type: CHANGE_COUNT_BY
value: 2 # Add 2 containers at a time
metric:
metricType: CPU_UTILIZATION
threshold:
operator: GT # Greater Than
value: 70 # 70% CPU usage
# Scale DOWN rules:
# If CPU is below 30% → remove 1 container instance
- displayName: scale-down-on-low-cpu
action:
type: CHANGE_COUNT_BY
value: -1 # Remove 1 container at a time
metric:
metricType: CPU_UTILIZATION
threshold:
operator: LT # Less Than
value: 30 # Below 30% CPU = we have too many containers
# Boundaries — never go below 1 or above 10 containers
capacity:
initialCount: 1
minCount: 1 # Always keep at least 1 running
maxCount: 10 # Never exceed 10 (cost control)
Always set a maxCount limit on autoscaling. Without it, a traffic spike (or a bug causing infinite retries) could spin up thousands of containers and give you a shock bill at month end! 💸
🔄 Versioning Models and CI/CD Pipelines for AI
Why Do We Need Versioning?
Imagine your team releases SmartHR version 2.0 with a new feature. But there is a bug — it gives wrong answers about sick leave. 10,000 employees are now reading wrong information. 😱
Without versioning, you cannot easily go back to version 1.0. With proper versioning — you roll back in 30 seconds. 🔄
Versioning is like the Undo button in Microsoft Word. 📝 Every save is a version. Something went wrong? Press Undo. Back to the good version instantly. CI/CD is the system that automatically saves, tests, and publishes every change.
SMARTHR VERSIONING STRATEGY: CONTAINER IMAGE TAGS: ───────────────────── iad.ocir.io/namespace/smarthr/chat-assistant:1.0.0 ← Initial release iad.ocir.io/namespace/smarthr/chat-assistant:1.1.0 ← Bug fix iad.ocir.io/namespace/smarthr/chat-assistant:2.0.0 ← New feature (risky!) iad.ocir.io/namespace/smarthr/chat-assistant:latest ← Always points to latest OCI FUNCTION VERSIONS: ────────────────────── smarthr-assistant (v1) ← Stable, 90% of traffic smarthr-assistant (v2) ← New version, 10% of traffic (Canary deployment) If v2 looks good after 1 day ──► Promote to 100% traffic If v2 has bugs ──► Roll back v1 to 100% traffic instantly CANARY DEPLOYMENT TRAFFIC SPLIT: ───────────────────────────────── Day 1: v1: 90% v2: 10% ← Watch for errors Day 2: v1: 70% v2: 30% ← Looking good, increase Day 3: v1: 50% v2: 50% ← Halfway there Day 4: v1: 0% v2: 100% ← Full deployment ✅
🏗️ CI/CD Pipeline with OCI DevOps
CI/CD stands for Continuous Integration / Continuous Deployment. It is the automatic system that takes your code changes — tests them, builds them, and deploys them — without anyone pressing a button.
SMARTHR CI/CD PIPELINE ON OCI DEVOPS:
Developer pushes code change
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ STAGE 1: BUILD (OCI DevOps Build Pipeline) │
│ ────────────────────────────────────────── │
│ • Run unit tests (pytest) │
│ • Check code quality (flake8, pylint) │
│ • Build Docker image │
│ • Push to OCI Container Registry with version tag │
│ (pass/fail in ~3 min) │
└────────────────────────┬────────────────────────────────────────┘
│
▼ (Only if all tests pass ✅)
┌─────────────────────────────────────────────────────────────────┐
│ STAGE 2: DEPLOY TO STAGING (Automatic) │
│ ────────────────────────────────────── │
│ • Deploy new version to staging environment │
│ • Run integration tests against staging │
│ • Run AI quality tests (check answer accuracy) │
│ (pass/fail in ~5 min) │
└────────────────────────┬────────────────────────────────────────┘
│
▼ (Only if staging tests pass ✅)
┌─────────────────────────────────────────────────────────────────┐
│ STAGE 3: CANARY DEPLOY TO PRODUCTION (10% traffic) │
│ ─────────────────────────────────────────────────── │
│ • Deploy to production — 10% of real users get new version │
│ • Monitor error rates and response times for 24 hours │
│ • If errors spike → automatic rollback to previous version 🔄 │
└────────────────────────┬────────────────────────────────────────┘
│
▼ (Only if canary looks healthy ✅)
┌─────────────────────────────────────────────────────────────────┐
│ STAGE 4: FULL PRODUCTION DEPLOY (100% traffic) │
│ ────────────────────────────────────────────── │
│ • Gradual ramp: 10% → 30% → 70% → 100% │
│ • Old version kept ready for 7 days (emergency rollback) │
│ • 🎉 All done! New SmartHR version is live! │
└─────────────────────────────────────────────────────────────────┘
This file is the instruction manual for OCI DevOps Build Pipeline. It tells OCI exactly what to do with every code push — run tests, build the image, and tag it with the version number. 🏷️
# ── FILE: build_spec.yaml ──────────────────────────────────────
# PURPOSE: Instructions for OCI DevOps Build Pipeline.
# Every time code is pushed, OCI follows these steps
# automatically — test, build, tag, done!
# ───────────────────────────────────────────────────────────────
version: 0.1
component: build
timeoutInSeconds: 1000
shell: bash
steps:
# Step 1: Run all unit tests FIRST
# If any test fails — stop everything. Do not build. Do not deploy.
- type: Command
name: Run Unit Tests
command: |
pip install -r requirements.txt
python -m pytest tests/ -v --tb=short
echo "✅ All tests passed!"
# Step 2: Check code quality
- type: Command
name: Code Quality Check
command: |
pip install flake8
flake8 func.py pdf_processor.py --max-line-length=100
echo "✅ Code quality check passed!"
# Step 3: Run AI Quality Tests
# Test that the AI gives correct answers to known questions
- type: Command
name: AI Answer Quality Test
command: |
python tests/test_ai_quality.py
# This file asks 20 test questions and checks the answers
# for accuracy, relevance, and safety
echo "✅ AI quality tests passed!"
# Step 4: Build the Docker image
- type: Command
name: Build Docker Image
command: |
# Use git commit hash as part of the version tag
export VERSION=$(cat version.txt)
export IMAGE_TAG="iad.ocir.io/namespace/smarthr/chat-assistant:${VERSION}"
docker build -t ${IMAGE_TAG} -t "iad.ocir.io/namespace/smarthr/chat-assistant:latest" .
echo "✅ Image built: ${IMAGE_TAG}"
outputArtifacts:
- name: smarthr_docker_image
type: DOCKER_IMAGE
location: "iad.ocir.io/namespace/smarthr/chat-assistant:${VERSION}"
🏢 Multi-Tenant Enterprise Access — Serving Many Companies at Once
What is Multi-Tenancy?
SmartHR becomes popular. Now five different companies want to use it: TechCorp, MediCare, LegalEagle, FinanceFirst, and EduLearn.
Each company has different HR policies, different employees, and different security needs. TechCorp employees should never see MediCare's private HR data. 🔒
Multi-tenancy means one deployment of SmartHR serves all five companies — but each company is completely isolated from the others. It is like one apartment building 🏢 with five separate apartments. One building, five private homes.
SMARTHR MULTI-TENANT ARCHITECTURE:
┌──────────────────────────────────────────────────────────────────────┐
│ OCI API GATEWAY │
│ (One entry point for all companies) │
└───────────────────────────┬──────────────────────────────────────────┘
│
JWT Token decoded → "tenant_id": "techcorp"
│
┌───────────────┼──────────────────┐
│ │ │
▼ ▼ ▼
┌──────────────────┐ ┌──────────────────┐ ┌──────────────────┐
│ TechCorp Space │ │ MediCare Space │ │ LegalEagle Space│
│ ─────────────── │ │ ─────────────── │ │ ────────────── │
│ OCI Compartment │ │ OCI Compartment │ │ OCI Compartment │
│ Oracle 23ai DB │ │ Oracle 23ai DB │ │ Oracle 23ai DB │
│ (their data │ │ (their data │ │ (their data │
│ only) │ │ only) │ │ only) │
└──────────────────┘ └──────────────────┘ └──────────────────┘
│ │ │
└──────────────────────┴──────────────────────┘
│
▼
┌────────────────────┐
│ OCI GenAI Service │
│ (Shared LLM — │
│ no data stored │
│ between calls) │
└────────────────────┘
🔑 How Does the Code Know Which Tenant to Use?
When an employee from TechCorp sends a question, the JWT token in their request contains their company name ("tenant_id": "techcorp"). This code reads that ID and only searches TechCorp's private Oracle 23ai data. It is impossible for TechCorp to accidentally see MediCare's data. 🔐
# ── FILE: multi_tenant_router.py ───────────────────────────────
# PURPOSE: Read the tenant_id from the user's JWT token and
# route the request to the correct company's data space.
# This ensures complete data isolation between companies.
# ───────────────────────────────────────────────────────────────
import jwt
import oracledb
import array
from fastapi import FastAPI, Request, HTTPException
from pydantic import BaseModel
app = FastAPI(title="SmartHR Multi-Tenant Router")
# Configuration per tenant
# Each company has their own Oracle 23ai schema (database user)
TENANT_CONFIG = {
"techcorp": {
"db_schema" : "TECHCORP_HR", # TechCorp's own schema
"compartment_id" : "ocid1.compartment.oc1..techcorp_xxxxx",
"display_name" : "TechCorp SmartHR"
},
"medicare": {
"db_schema" : "MEDICARE_HR", # MediCare's own schema
"compartment_id" : "ocid1.compartment.oc1..medicare_xxxxx",
"display_name" : "MediCare SmartHR"
},
"legaleagle": {
"db_schema" : "LEGALEAGLE_HR",
"compartment_id" : "ocid1.compartment.oc1..legal_xxxxx",
"display_name" : "LegalEagle SmartHR"
}
}
class QuestionRequest(BaseModel):
question: str
def extract_tenant_id(request: Request) -> str:
"""
Read the JWT token from the Authorization header.
Decode it to find which company this user belongs to.
If no valid token → reject the request with 401 Unauthorized.
"""
auth_header = request.headers.get("Authorization", "")
if not auth_header.startswith("Bearer "):
raise HTTPException(status_code=401, detail="No valid token provided")
token = auth_header.replace("Bearer ", "")
try:
# Decode the JWT (OCI Identity validates the signature)
payload = jwt.decode(token, options={"verify_signature": False})
tenant_id = payload.get("tenant_id", "").lower()
if tenant_id not in TENANT_CONFIG:
raise HTTPException(status_code=403, detail=f"Unknown tenant: {tenant_id}")
return tenant_id
except Exception as e:
raise HTTPException(status_code=401, detail=f"Invalid token: {str(e)}")
@app.post("/ask")
async def ask_question(request: Request, body: QuestionRequest):
"""
Main endpoint — reads tenant from JWT, routes to correct company data.
"""
# Step 1: Find out which company this user is from
tenant_id = extract_tenant_id(request)
tenant_cfg = TENANT_CONFIG[tenant_id]
# Step 2: Only query THIS company's data in Oracle 23ai
# The DB schema (e.g., TECHCORP_HR) ensures complete isolation
chunks = search_tenant_data(
question = body.question,
db_schema = tenant_cfg["db_schema"],
top_k = 3
)
# Step 3: Generate answer using shared OCI GenAI service
# (LLM is shared, but no tenant data is stored in the LLM)
answer = generate_answer(body.question, chunks)
return {
"tenant" : tenant_cfg["display_name"],
"question": body.question,
"answer" : answer
}
def search_tenant_data(question: str, db_schema: str, top_k: int) -> list:
"""
Search ONLY the specified tenant's Oracle 23ai schema.
TechCorp users ONLY search TECHCORP_HR schema.
MediCare users ONLY search MEDICARE_HR schema.
"""
# ... (embedding + Oracle 23ai VECTOR_DISTANCE search)
# Uses db_schema to prefix the table: e.g., TECHCORP_HR.HR_KNOWLEDGE
pass
- 🗄️ Database level: Give each tenant their own Oracle 23ai schema or database user
- ☁️ Compartment level: Use separate OCI compartments per enterprise tenant
- 🔑 IAM level: Each tenant has their own OCI IAM group and policies
- 📊 Logging level: Separate OCI Logging streams per tenant for audit trails
- 💰 Cost level: Use OCI Cost Tracking tags to bill each tenant separately
Never use a single database table with a
tenant_id column as the only isolation.
A bug in a WHERE clause could leak one company's data to another.
Always use schema-level or database-level isolation as the primary boundary. 🔐
📊 Monitoring Your Deployed SmartHR App
Deploying is not the end. You need to watch your app like a pilot watches an airplane dashboard. 🛩️ If something goes wrong — you want to know in seconds, not hours.
SMARTHR MONITORING DASHBOARD — KEY METRICS TO WATCH: ┌──────────────────────────────────────────────────────────────────────┐ │ │ │ 📊 METRIC 1: Response Time │ │ Target: < 3 seconds for every answer │ │ Alert: > 5 seconds = something is slow. Investigate! │ │ │ │ 📊 METRIC 2: Error Rate │ │ Target: < 0.1% of requests fail │ │ Alert: > 1% errors = something is broken. Page the on-call team! │ │ │ │ 📊 METRIC 3: OCI GenAI Token Usage │ │ Target: Monitor tokens/minute to stay within service limits │ │ Alert: > 80% of quota = scale up or add rate limiting │ │ │ │ 📊 METRIC 4: Concurrent Function Executions │ │ Target: Track peak concurrency │ │ Alert: Approaching concurrency limit = scale up proactively │ │ │ │ 📊 METRIC 5: Answer Quality Score (AI-powered monitoring!) │ │ Target: > 90% of sampled answers rated "accurate" by a judge LLM │ │ Alert: < 80% accuracy = model quality degraded. Roll back! │ │ │ └──────────────────────────────────────────────────────────────────────┘
This code sends custom metrics from the SmartHR app to OCI Monitoring. It records how long each AI answer took, and whether it succeeded or failed. OCI Monitoring stores these numbers and can alert you if things go wrong. 📊
# ── FILE: monitoring.py ────────────────────────────────────────
# PURPOSE: Send SmartHR performance metrics to OCI Monitoring.
# Like a flight recorder for your AI app —
# tracks every request: how fast, did it succeed, which tenant.
# ───────────────────────────────────────────────────────────────
import oci
import time
from datetime import datetime, timezone
class SmartHRMonitor:
"""Sends real-time metrics to OCI Monitoring service."""
def __init__(self, compartment_id: str):
config = oci.config.from_file("~/.oci/config")
self.monitoring = oci.monitoring.MonitoringClient(config)
self.compartment_id = compartment_id
def record_request(self, tenant_id: str, duration_ms: float, success: bool):
"""
Record one SmartHR request to OCI Monitoring.
Called after every single API request completes.
"""
timestamp = datetime.now(timezone.utc)
# Build the metric data points
datapoints = [
oci.monitoring.models.Datapoint(
timestamp = timestamp,
value = duration_ms # How many milliseconds the request took
)
]
# Send "response time" metric to OCI Monitoring
self.monitoring.post_metric_data(
oci.monitoring.models.PostMetricDataDetails(
metric_data = [
oci.monitoring.models.MetricDataDetails(
namespace = "smarthr_app", # Our custom namespace
name = "response_time_ms", # Metric name
compartment_id = self.compartment_id,
dimensions = {
"tenant_id" : tenant_id, # Which company
"status" : "success" if success else "error"
},
datapoints = datapoints
)
]
)
)
# ── HOW TO USE IT IN THE FUNCTION ──────────────────────────────
monitor = SmartHRMonitor(compartment_id="ocid1.compartment.oc1..xxxxx")
def handler(ctx, data):
start_time = time.time()
success = True
tenant_id = "unknown"
try:
# ... (all the SmartHR logic here) ...
pass
except Exception as e:
success = False
raise e
finally:
# Always record the metric — success OR failure
duration_ms = (time.time() - start_time) * 1000
monitor.record_request(tenant_id, duration_ms, success)
🗺️ The Complete SmartHR Deployment Architecture
SMARTHR — COMPLETE OCI DEPLOYMENT ARCHITECTURE (2026) ┌──────────────────────────────────────────────────────────────────────────────┐ │ │ │ DEVELOPMENT CI/CD PRODUCTION │ │ ────────── ────────── ────────── │ │ │ │ 👩💻 Developer OCI DevOps OCI API Gateway │ │ pushes code ──────► Build Pipeline ──────► (Rate Limiting, │ │ to Git repo (Test → Build Auth, Routing) │ │ → Push image) │ │ │ │ │ │ ┌──────────┴──────────┐ │ │ │ │ │ │ ▼ ▼ │ │ OCI Functions OCI Container│ │ (Chat Q&A Instances │ │ serverless) (PDF Ingest) │ │ │ │ │ │ └──────────┬──────────┘ │ │ │ │ │ ▼ │ │ ┌───────────────────────────┐ │ │ │ Oracle Database 23ai │ │ │ │ (Separate schemas per │ │ │ │ tenant — full isolation)│ │ │ └──────────────┬────────────┘ │ │ │ │ │ ▼ │ │ ┌───────────────────────────┐ │ │ │ OCI Generative AI │ │ │ │ Cohere Command R+ │ │ │ │ / Llama 3.3 │ │ │ └───────────────────────────┘ │ │ │ │ MONITORING: OCI Monitoring + OCI Logging + OCI Alarms (24/7) 📊 │ │ SECURITY: OCI Vault + OCI IAM + OCI WAF (Web Application Firewall) 🔒 │ │ │ └──────────────────────────────────────────────────────────────────────────────┘
⚠️ Common Deployment Mistakes (And How to Avoid Them)
Always use: Development → Staging → Canary → Full Production. Skipping stages to save time is how you cause major incidents at 2 AM. 😴🚨
Use OCI Vault for all secrets. Never write passwords, API keys, or database credentials directly in code files. Anyone who reads your code would have full access to your cloud account. 🔓
A new LLM version might give different (worse) answers. Always run an AI Quality Test Suite (20–50 golden test questions) before promoting any new model version to production. 🧪
Always set a reasonable maximum. A traffic spike or a bug causing retries can spin up thousands of containers and create a bill shock. 💸
{"tenant": "techcorp", "env": "production", "version": "2.1.0"}.
Tags make monitoring, billing, and troubleshooting 10x easier. 🏷️
⚡ Quick Reference — OCI Deployment Cheat Sheet
- ⚡ OCI Functions — Serverless, auto-scales to zero, pay per call, max 30s timeout, best for quick AI Q&A
- 🐳 OCI Container Instances — Always-on containers, max 32 GB RAM, GPU support, best for heavy jobs
- 🌐 OCI API Gateway — Single front door, JWT auth, rate limiting, multi-tenant routing
- 📦 OCI Container Registry — Store and version your Docker images with semantic version tags
- 🔄 OCI DevOps — CI/CD pipelines, canary deployments, automated testing and rollback
- 📊 OCI Monitoring + Alarms — Custom metrics, real-time dashboards, alert on anomalies
- 🔒 OCI Vault — Store all secrets, never in code files
- 🏷️ OCI Resource Manager — Terraform-based infrastructure as code for reproducible deployments
Happy deploying! 🛠️✨
Comments
Post a Comment