Skip to main content

Deploying Deep Learning Models: Inference, APIs, Batching, and Quantization

Calculating read time…

Deploying a deep learning model means turning a trained set of weights into a service that can answer requests reliably, at a predictable cost, under real traffic — which is a completely different engineering problem than training the model in the first place. Inference serving, request batching, and quantization are the three levers that decide whether that service is fast, cheap, and stable, or slow, expensive, and fragile. 🧠

When a model serves millions of daily requests, small inefficiencies compound: an unbatched request wastes most of a GPU's compute, an unquantized model doubles your memory footprint and cloud bill, and a single-instance deployment becomes a single point of failure the moment traffic spikes or a host reboots. Getting inference architecture right is what separates a model that works in a demo from one that survives a product launch. ⚙️

Diagram showing a request flowing from a user through an API gateway and load balancer into a batching queue, processed by a GPU worker pool running a quantized model, with results cached and an autoscaler and observability layer monitoring queue depth, GPU utilization, and latency

🔀 Quick Comparison: Precision Formats for Inference

Format Relative memory Typical accuracy impact Best fit
FP32 Baseline (largest) None — reference precision Training, debugging, correctness baselines
FP16 / BF16 ~50% of FP32 Usually negligible Default production inference on GPUs
INT8 ~25% of FP32 Small, workload-dependent; needs calibration Latency- or cost-sensitive services at scale
INT4 / mixed ~12–15% of FP32 Noticeable on some tasks; needs strong evaluation Very large models where memory is the hard constraint

Numbers are directional engineering rules of thumb, not measurements from a specific benchmark; always validate on your own model and task before choosing a precision.

1. Foundations: What "Inference" Actually Means in Production

Child-friendly analogy: Imagine a bakery that spent months perfecting a cake recipe (that's training). Inference is what happens every single day after that: a customer walks in, orders a cake, and someone has to actually bake it, on demand, correctly, quickly, and without running out of flour — for every customer, all day, every day.

Technically, inference is the process of running a trained model's forward pass on new input to produce an output — a classification, a translation, a generated token sequence, an embedding, or a detection box. Training happens once (or periodically); inference happens continuously, driven by unpredictable real-world traffic. That single difference reshapes every engineering decision: training optimizes for eventual convergence, while inference optimizes for tail latency, throughput per dollar, and availability.

Three concepts sit at the center of any serious inference system:

  • Serving API — the contract (usually HTTP/REST, sometimes gRPC) that lets applications send input and receive predictions without knowing anything about GPUs, frameworks, or model internals.
  • Batching — grouping multiple requests so the accelerator processes them together, since GPUs are throughput machines that are inefficient when fed one small request at a time.
  • Quantization — reducing the numerical precision of a model's weights (and sometimes activations) so it needs less memory and often runs faster, at a cost some accuracy must be measured against.

💡 Trade-off to internalize early: Every technique in this post trades something for something else — batching trades a small amount of latency for a large amount of throughput; quantization trades a small amount of numerical precision for a large amount of memory and cost savings. There is no free lunch, only well-measured trade-offs.

🎯 Use this when: you are deciding whether a model even needs specialized inference engineering, or whether a simple always-on endpoint is enough for your traffic level.

2. Mechanics: The Request Lifecycle, Batching Strategies, and Quantization

2.1 The lifecycle of a single inference request

Analogy: Think of airport security. Your boarding pass gets checked (authentication), you're routed to whichever lane is shortest (load balancing), your bag might wait for others going through the same scanner (batching), and the scanner itself has been tuned to move fast without missing anything important (the optimized model).

  1. Ingress and authentication: the request hits an API gateway, which validates the API key or token, checks the request shape, and applies rate limits.
  2. Routing: a load balancer forwards the request to a healthy inference instance, based on health checks and current load.
  3. Queueing and batching: the serving layer holds the request briefly (microseconds to a few tens of milliseconds) to see if it can be grouped with others.
  4. Forward pass: the batched tensor runs through the model on the accelerator.
  5. Post-processing: raw logits or tensors are converted into the response format the caller expects (a label, JSON, generated text).
  6. Response and logging: the result is returned, and latency, token counts, and any errors are emitted as metrics for observability.

Every one of these steps can fail independently, which is why production systems need health checks, timeouts, and retries at each hop rather than treating "call the model" as one atomic operation.

2.2 Batching strategies

There isn't one kind of batching — the right strategy depends on whether requests are uniform in size and whether latency budgets are tight:

  • Static batching: the server waits until it has a fixed number of requests (or a timeout expires), then runs them together. Simple, but a single slow request in the batch delays everyone else in it, and empty slots waste compute.
  • Dynamic batching: the server continuously forms batches from whatever requests have arrived within a short window, adapting batch size to current load. This is the most common general-purpose approach for classification, embedding, and vision models.
  • Continuous (in-flight) batching: used for autoregressive text generation, where requests generate output token-by-token at different speeds. Rather than waiting for the slowest sequence in a batch to finish, the scheduler adds and removes individual sequences from the running batch as they complete, keeping the GPU busy. This is the mechanism that made high-throughput LLM serving practical, and it underlies inference engines such as vLLM.

✅ Worked example: A recommendation service receiving 200 requests/second, each needing ~8ms of raw GPU compute if run alone, would waste most of the GPU's parallel capacity processing them one at a time. Grouping requests that arrive within a 5–10ms window into batches of 16–32 can let the same GPU serve several times the traffic, because the accelerator's matrix units are being kept busy on a larger structured operation instead of a run of tiny ones.

2.3 Quantization mechanics

Analogy: Imagine measuring a room with a tape measure marked in millimeters versus one marked only in centimeters. The coarser tape is faster to read and lighter to carry, and for most everyday purposes — hanging a picture, buying a rug — the lost precision doesn't matter. But if you're fitting custom cabinetry, that same coarseness can throw the whole job off. Quantization is choosing a coarser "ruler" for a model's numbers.

Neural network weights and activations are normally stored as 32-bit floating point (FP32) numbers during training. Quantization represents those numbers with fewer bits — commonly 16-bit floating point (FP16/BF16), 8-bit integers (INT8), or, increasingly, 4-bit formats for very large models. Two main approaches exist:

  1. Post-training quantization (PTQ): take an already-trained model and convert its weights to lower precision, typically using a small calibration dataset to decide how to map the range of real values onto the reduced set of representable values. Fast to apply, but can lose more accuracy on sensitive models.
  2. Quantization-aware training (QAT): simulate the lower precision during training or fine-tuning, so the model learns weights that are robust to the eventual rounding. More effort, but typically preserves accuracy better, especially at very low bit widths.

The benefit is threefold: a smaller model fits in less GPU memory (letting you either use a smaller, cheaper GPU or fit more model replicas on the same GPU), memory bandwidth — often the real bottleneck in modern inference — is reduced, and on hardware with native low-precision support, raw compute throughput can also increase. The cost is a measurable, task-dependent accuracy shift that must be evaluated, never assumed.

🎯 Use this when: you understand why "just run it on a bigger GPU" is usually a worse first move than batching and quantizing the workload you already have.

3. A Worked Example: Serving an LLM on OCI Data Science Model Deployment

OCI's managed path for this workload is the OCI Data Science Model Deployment service. According to Oracle's documentation, it exposes a trained or fine-tuned model as a managed HTTPS endpoint, and its AI Quick Actions feature supports deploying foundation models directly with inference frameworks such as vLLM or TGI, with the resulting endpoint mimicking the OpenAI chat completion API by default.[1][2] vLLM is a widely used open-source engine specifically built around continuous batching and efficient GPU memory management (via its "PagedAttention" memory allocator) for LLM serving — which is exactly the batching mechanism described in section 2.2.

Oracle's documentation also describes a metric-based autoscaling capability for Model Deployments: you define minimum, initial, and maximum instance counts, plus scale-in/scale-out thresholds evaluated against a chosen metric (built-in or a custom Monitoring Query Language expression), and the service can optionally scale the attached load balancer's bandwidth within a configured range at the same time. A cool-down period after each scaling action prevents thrashing.[3][4]

✅ Hypothetical, labeled example: A customer-support summarization feature fine-tunes an open-weight LLM, quantizes it to FP16 for a good latency/quality balance, and deploys it via AI Quick Actions with vLLM as the serving engine on a Model Deployment backed by a single-GPU shape. Autoscaling is configured to add instances when GPU utilization or queue-wait-time crosses a threshold, and the load balancer's bandwidth range is set to grow alongside instance count during a product launch. This is an illustrative architecture, not a documented case study.

For GPU capacity, Oracle publishes a range of compute shapes for Data Science Model Deployment, from single-GPU VM shapes (for example, VM.GPU.A10.1) suited to smaller models and lower-traffic endpoints, up to multi-GPU bare metal shapes such as BM.GPU.A100-v2.8 or BM.GPU.H100.8 for large models or high-throughput serving.[5] Bare metal GPU shapes include NVLink for GPU-to-GPU bandwidth within the node, while VM GPU shapes do not — a distinction that matters if you plan to shard a very large model across GPUs with tensor parallelism rather than replicate a smaller one.[6]

🎯 Use this when: you're choosing between building your own serving stack on general-purpose compute versus using a managed model-deployment service that already implements health checks, versioned endpoints, and autoscaling.

4. Implementation: Building the OCI Inference Architecture

A production inference stack on OCI is rarely one service — it's a small number of services composed with clear responsibilities:

  1. API Gateway: terminates public traffic, enforces authentication, request-size limits, and coarse rate limiting before anything reaches the model.
  2. Compute layer: OCI Data Science Model Deployment for a managed, framework-aware serving endpoint; OCI Container Engine for Kubernetes (OKE) with the NVIDIA GPU Operator when you need custom serving containers (for example, NVIDIA Triton Inference Server or NVIDIA NIM microservices) and finer control over scheduling; or OCI Container Instances for simple, short-lived, or bursty inference containers that don't need full Kubernetes.
  3. GPU shape selection: match the shape family to the workload — single-GPU VM shapes for smaller models or development, multi-GPU bare metal shapes with NVLink for large models needing tensor parallelism, and NVIDIA GPU Operator integration on OKE for driver and device-plugin management.[6]
  4. Secrets and networking: store API keys, database credentials, and model registry credentials in OCI Vault rather than in code or container images; place inference instances in private subnets with security lists or network security groups restricting inbound traffic to the load balancer only.
  5. Artifact and image management: store versioned model artifacts in OCI Object Storage and container images in OCI Registry (OCIR), so every deployment references an immutable, reproducible artifact rather than "whatever is on the box."
  6. Autoscaling: for Model Deployments, configure metric-based compute and load-balancer bandwidth autoscaling as described in section 3; for OKE, pair the Kubernetes Horizontal Pod Autoscaler (for replica count) with the Cluster Autoscaler (for node/GPU pool sizing), driven by GPU utilization or queue-depth metrics rather than CPU alone, since CPU is rarely the bottleneck in GPU inference.
  7. Observability: emit latency percentiles (not just averages), request/error rates, GPU utilization and memory, batch size distribution, and token throughput (for generative models) to OCI Monitoring and Logging, with dashboards and alarms tied to the same thresholds used for autoscaling and incident response.
# Illustrative only — not copied from any vendor example.
# Conceptual shape of a Model Deployment autoscaling rule.
scaling_metric: "GPU_UTILIZATION"
scale_out_threshold_pct: 70
scale_in_threshold_pct: 25
min_instances: 2
max_instances: 12
cooldown_seconds: 300
load_balancer_bandwidth_range_mbps: [10, 20]

💡 Trade-off: Oracle's guidance notes that the maximum load-balancer bandwidth in an autoscaling configuration can be set to no more than twice the minimum — a documented ceiling worth designing around rather than discovering during a launch.[3] Plan your minimum bandwidth for expected steady-state load, not just for the smallest traffic you've ever seen.

Minimum instance counts of one or two should generally be avoided for anything user-facing: with a minimum of one, a single instance restart or fault-domain issue means a full outage; with two spread across fault domains, the service degrades rather than disappears while the autoscaler or platform replaces the unhealthy instance. Multi-availability-domain placement, where the region supports it, further reduces the blast radius of a single data-center-level event; note that not every OCI region has multiple availability domains, so region choice affects how much of this redundancy is available out of the box, and this should be verified against current OCI region documentation for your target region rather than assumed.

🎯 Use this when: you're sketching the actual service boundaries and OCI resources for a new inference workload, not just picking an algorithm.

5. Step-by-Step: Deploying Your First Model on OCI

Everything above explains why the pieces exist. This section is the beginner's walkthrough for actually standing one up. The fastest realistic path uses OCI Data Science AI Quick Actions, a low-code feature built specifically so a first deployment doesn't require writing a serving container.[2][7]

5.1 Path A — Deploy a catalog model with AI Quick Actions (no code)

  1. Set up IAM policies. Grant your user group permission to manage Data Science projects, notebooks, and model deployments in the target compartment. Missing this step is the most common reason a first attempt fails silently.
  2. Request GPU quota if needed. Under Governance & Administration → Limits, Quotas and Usage, check your GPU shape limit and file a service-limit-increase request if it is at zero. Approval is not instant, so do this before you plan to deploy.
  3. Create a Data Science Project and Notebook Session. A small CPU shape is enough just to open the notebook UI — you don't need a GPU for this step.
  4. Open AI Quick Actions from the notebook. It appears under Extensions in the Notebook Launcher, and gives access to Models, Deployments, and Evaluations.[7]
  5. Choose a model and create a deployment. Pick a foundation model tagged "Ready to Deploy" in the Model Explorer (or a fine-tuned model of your own), select a GPU compute shape sized to that model, and optionally enable logging — recommended for troubleshooting.[8]
  6. Set instance count and load balancer bandwidth under advanced options. Start with a single instance to validate the deployment end to end before adding redundancy.
  7. Wait for the deployment to reach "Active," then call the generated HTTPS endpoint. By default the endpoint mimics the OpenAI chat completion API, so existing OpenAI-compatible client code usually works with minimal changes.[2]

✅ Practical tip: Prove the deployment works end to end with one instance and no autoscaling first. Only after a successful test call should you layer on multiple instances and the autoscaling rules from section 4 — debugging autoscaling and a broken endpoint at the same time is much harder than debugging them one at a time.

5.2 Path B — Bring your own fine-tuned model

  1. Upload your model artifact to a versioned OCI Object Storage bucket. Versioning matters because it gives you a rollback point if a later artifact is bad.[9]
  2. Grant the Data Science service read (and write) access to that bucket through an IAM policy.
  3. Register the model inside AI Quick Actions and deploy it the same way as Path A, step 5 onward, selecting your registered model instead of a catalog one.[9]

5.3 Path C — Full manual deployment (advanced)

Reach for this only once Path A works and you need something it doesn't offer, such as a custom serving container, a non-LLM model type, or a specific framework version pinned outside AI Quick Actions' supported set.

  1. Package a model artifact following the Data Science Model Deployment artifact format and register it in the Model Catalog.
  2. Create the deployment manually — via the console's Create Model Deployment flow or the oci data-science model-deployment create CLI command — pointing at that catalog model.
  3. Configure autoscaling explicitly using the parameters covered in section 4 (scaling metric, min/max instances, load balancer bandwidth range, cool-down period).
  4. Attach logging and call the resulting endpoint the same way as Path A.

💡 Beginner traps: the three most common first-deployment failures are skipping the IAM policy step, forgetting that new tenancies start with zero GPU quota, and turning off logging to "keep things simple" — which then leaves nothing to debug a failed deployment with.

🎯 Use this when: you're deploying a model on OCI for the first time and need the actual click-by-click and CLI path, not just the architecture behind it.

6. Enterprise Rollout: Governance, Evaluation, and Cost Control

Shipping a model once is easy; operating it responsibly for millions of daily users, indefinitely, is the actual job. A few practices consistently separate durable deployments from fragile ones:

5.1 Ownership and governance

Every deployed model version needs a named owner, a documented rollback procedure, and access controls (IAM policies scoped to least privilege) governing who can push a new version to production. Treat the model artifact, the serving configuration, and the infrastructure-as-code definitions as one versioned unit — deploying a new model weight file against stale serving code (or vice versa) is a common source of silent regressions.

5.2 Evaluation before and after deployment

  • Offline evaluation on a versioned, representative held-out test set before any release — with strict separation from training data to avoid leakage.
  • Quantization-specific regression testing: re-run the same evaluation suite on the quantized model, not just the full-precision one, since accuracy drift from quantization can be uneven across subgroups or task types.
  • Canary releases: route a small percentage of real traffic to the new model version and compare latency, error rate, and quality metrics against the current production version before a full rollout.
  • Online monitoring for drift: track whether the distribution of production inputs or outputs is shifting away from what the model was evaluated on, since accuracy silently decays as the world changes even when the model itself hasn't.
  • Human review sampling for generative outputs, particularly for high-stakes or customer-facing text, alongside automated metrics.
  • Rollback criteria defined in advance: a canary should have explicit, pre-agreed thresholds (for example, error-rate or latency regression beyond a set margin) that trigger an automatic or manual rollback, rather than a judgment call made under pressure during an incident.

5.3 Cost and capacity discipline

GPU capacity is typically the largest line item in an inference budget. Set budget alerts and cost dashboards tied to the compartments running inference workloads, cap maximum autoscaling instance counts so a traffic anomaly or bug can't scale spend unboundedly, and periodically re-evaluate whether a cheaper GPU shape or a more aggressive quantization level still meets your quality bar as the model or traffic pattern changes.

5.4 Incident response

Write runbooks before you need them: what to check first when p95 latency spikes (queue depth and batch size are usually the first place to look before assuming a hardware fault), how to roll back a model version, and who is paged when GPU utilization pins at 100% with growing queue depth — a symptom of insufficient capacity rather than a code bug.

🎯 Use this when: your model is moving from "a working prototype" to "a system other teams and customers depend on."

7. Common Mistakes (and Why They Hurt)

Deploying without batching, then blaming the GPU. A common pattern is to serve one request per forward pass and conclude the GPU is "too slow," when the real issue is that the accelerator's parallel compute is barely used at batch size one. The fix is architectural (add a batching layer or use a serving engine that already implements it), not a bigger GPU, which would simply repeat the same underutilization at higher cost.

Quantizing without re-evaluating. Teams sometimes quantize a model, see that it "still works" on a handful of manual spot checks, and ship it. Because quantization error is rarely uniform, this can hide regressions on specific input types, languages, or edge cases that only show up at scale — precisely the failure a proper offline evaluation suite is meant to catch before it reaches production traffic.

Setting a minimum instance count of one. This looks fine in testing and fails the first time that single instance needs to restart, get patched, or hits a hardware fault — at which point the service has zero capacity rather than degraded capacity. Two or more instances across fault domains is the minimum viable pattern for anything user-facing.

Autoscaling on the wrong metric. Scaling GPU inference workloads on CPU utilization alone frequently fails to trigger scale-out, because the GPU — not the CPU — is usually the bottleneck; the service degrades in latency while the autoscaler sees "normal" CPU and does nothing.

No rollback plan for a bad model version. Without a versioned artifact and a tested rollback path, a bad deployment turns a five-minute fix into an extended incident, because the team has to reconstruct — under pressure — how to get back to the last known-good state.

Treating average latency as the health signal. Averages hide the experience of the slowest fraction of users. A system can have a fine average latency while 5% of requests time out, because batching, GPU contention, or a noisy neighbor instance creates a long tail that averages wash out. Track p95/p99 latency, not just the mean.

❓ FAQ

Does batching always improve performance?

No — batching improves throughput (requests handled per second) at the cost of some added per-request latency, since a request may wait briefly to be grouped with others. For latency-critical, low-traffic endpoints, a very small batch window or no batching may be the better trade-off.

Is quantization safe for every model?

Not automatically. Sensitivity to reduced precision varies by architecture, task, and even by specific layers within a model. Always run your full evaluation suite — including subgroup and edge-case checks — on the quantized version before deploying it, rather than assuming accuracy is preserved.

Should I use OCI Data Science Model Deployment or build my own serving stack on OKE?

Model Deployment is generally the faster path when your model fits its supported frameworks (including vLLM- or TGI-based serving via AI Quick Actions) and you want managed autoscaling and endpoints out of the box. OKE with the NVIDIA GPU Operator makes sense when you need custom serving containers, non-standard networking, or tighter control over scheduling across a larger, multi-workload cluster.

How many GPU instances do I need for high availability?

There's no universal number, but a minimum of two instances spread across fault domains is a reasonable floor for user-facing services, so a single instance failure degrades rather than eliminates capacity. Actual sizing should come from load testing against your real traffic and latency targets.

Can this article guarantee my deployment will never go down?

No. No cloud architecture — on OCI or any provider — can honestly promise zero downtime or infinite scale. The realistic goal is high availability through redundancy, tested rollback, capacity headroom, and graceful degradation, which minimizes user-visible impact rather than eliminating all risk.

🔗 References & Further Reading

  1. [1] Oracle, "About Model Deployment," OCI Data Science documentation — docs.oracle.com
  2. [2] LangChain, "ChatOCIModelDeployment integration" (describing AI Quick Actions, vLLM/TGI support, and the OpenAI-compatible endpoint) — docs.langchain.com
  3. [3] Oracle, "Model Deployment Autoscaling," OCI Data Science documentation — docs.oracle.com
  4. [4] Oracle, "Creating a Model Deployment with Autoscaling," OCI Data Science documentation — docs.oracle.com
  5. [5] Oracle, "Supported Compute Shapes," OCI Data Science documentation — docs.oracle.com
  6. [6] Oracle, "New OCI Compute Bare Metal and VM GPU Shapes," OCI blog — blogs.oracle.com
  7. [7] Oracle, "AI Quick Actions in OCI Data Science," OCI blog — blogs.oracle.com
  8. [8] Oracle, "Model Deployment," AI Quick Actions documentation — docs.oracle.com
  9. [9] Oracle, "AI Quick Actions: Bring Your Own Model," OCI blog — blogs.oracle.com

📝 Summary

  • Inference is a continuous, latency- and cost-sensitive operational problem, distinct from training.
  • Batching (static, dynamic, or continuous) trades a little latency for much higher GPU throughput.
  • Quantization trades a measured amount of accuracy for large memory and cost savings, and must always be re-evaluated, not assumed.
  • OCI Data Science Model Deployment provides a managed path with autoscaling for compute and load-balancer bandwidth; OKE offers more control for custom serving stacks.
  • A first deployment is fastest through AI Quick Actions: set IAM policies and GPU quota, launch a notebook, then deploy a catalog or fine-tuned model to a working endpoint with a few clicks.
  • Enterprise rollout depends on ownership, versioned evaluation, canaries with pre-agreed rollback criteria, drift monitoring, and cost guardrails.
  • Most production incidents in this space trace back to a handful of avoidable mistakes: no batching, unvalidated quantization, single-instance deployments, wrong autoscaling metrics, and no rollback plan.

Thanks for reading — go batch something. 🚀

Comments