---
title: "How to Deploy a Fine-Tuned Model as an API"
slug: deploy-fine-tuned-model-as-api
description: "Your fine-tuned model works in the notebook. Here's how to serve it as a production API with auth, autoscaling and logs, and how to do it on TIR."
author: "Hemant Singh Tanwar"
date: "September 29, 2026"
category: "AI & Machine Learning"
faqAccordion: true
featured_image: "https://strapi-main-website.objectstore.e2enetworks.net/deploy_fine_tuned_model_as_api_5c04d547b5.png"
---

To deploy a fine-tuned model as an API, you need four things the notebook never asked for: somewhere to keep the weights, a serving runtime that fits the architecture, an authenticated HTTPS endpoint, and scaling rules so you are not paying for idle GPUs. On [TIR Model Endpoints](https://www.e2enetworks.com/inference-endpoints) that is three choices in a console — point at the model, pick a runtime and GPU, set scaling and generate a token — and you get an OpenAI-compatible URL with logs, metrics and alerts already wired in.

Here is the same thing as a picture, and then the detail behind each step.

![TIR Model Endpoints in three steps: bring your model from Hugging Face, GitHub or the TIR Model Repository; pick how to serve it with vLLM, SGLang or a custom container; deploy with auth tokens, autoscaling, and logs, metrics and alerts](https://strapi-main-website.objectstore.e2enetworks.net/tir_model_endpoints_3_steps_1_fdbfb10601.png)

## Why models stall between training and production

Your fine-tuned model scores well on the eval set. The notebook runs clean. Then someone from the app team asks a fair question: "Can we start calling it by Friday?"

That question is where a lot of good models stop moving. Training is the part everyone planned for. Turning the model into something another team can call over HTTP, safely and reliably, is the part that shows up afterwards.

The work usually falls between teams. The ML engineers built the model and would rather keep improving it. The platform or DevOps team runs infrastructure but may not know the serving stack well. The app team just wants a URL.

So the model waits. Someone sets up a quick server on a VM to unblock a demo, and that server quietly becomes production, with no autoscaling and logs that live on one machine. Or the project slips a sprint while people work out who owns it.

The fix is less about effort and more about removing the step altogether: if deploying a model is as routine as deploying any other service, it stops needing an owner.

## What "production-ready" means for a model endpoint

A model in a notebook answers one person. A model behind an endpoint has to answer whoever calls it, at whatever hour, without someone watching the terminal. That brings in a handful of requirements that have nothing to do with the model itself.

- **A stable API with access control.** Other services need a URL and a way to prove they are allowed to use it. That means tokens, and a way to issue and revoke them.
- **A serving runtime that suits the model.** How you load the weights and batch requests affects latency and how many requests one GPU can handle. The runtime you pick matters more than people expect.
- **Scaling in both directions.** Traffic is rarely flat. You want more replicas when requests pile up, and fewer (or none) when things go quiet, so you are not paying for idle GPUs overnight.
- **Visibility when things go wrong.** When a request fails at 11 pm, someone needs logs to see why, metrics to spot a slowdown, and alerts so they hear about it before a customer does.
- **A sensible network boundary.** Many teams want inference traffic to stay on a private network rather than cross the public internet, especially when prompts carry customer data.

None of these are hard problems on their own. Together, they are a small infrastructure project.

## Merged weights or LoRA adapter: what do you actually deploy?

If you fine-tuned with LoRA or QLoRA — which most teams do — your training run produced an adapter, not a whole model. A [fine-tuning job on TIR](https://www.e2enetworks.com/training-fine-tuning) writes every checkpoint and adapter to your TIR Model Repository, so you arrive at deployment with a decision to make.

Both paths work, and the choice is mostly about how many fine-tunes you are serving.

**Merge the adapter into the base weights, then deploy the merged model.** This is the common path and the one to default to. You get a single self-contained model, any runtime can serve it, and there is no adapter-loading overhead at inference time. The cost is disk: each fine-tune is a full copy of the base model.

**Serve the adapter on top of the base model.** Worth it when you have several fine-tunes of the same base and want them behind one deployment rather than one GPU each. You keep one copy of the base weights and swap adapters per request, at the cost of a little latency and a runtime that supports it.

If you are serving one fine-tune, merge it. The disk is cheaper than the complexity.

## How to deploy a fine-tuned model on TIR in three steps

Model Endpoints on TIR are built so that deploying a model is a few choices in a console rather than a project. The flow has three steps.

**1. Bring your model.** Point the endpoint at a model on Hugging Face, in a GitHub project, or in your TIR Model Repository (backed by E2E Object Storage). If your setup is unusual, you can bring your own Docker container from a public or private registry instead.

**2. Pick how to serve it.** Choose a runtime, such as vLLM, SGLang or Triton, or run your own container. Then choose the GPU and how many workers you want.

**3. Deploy.** Set your scaling rules and create an API token. The endpoint comes with autoscaling, centralised and per-replica logs, hardware and service metrics, access logs and email alerts already set up. Before you share the URL with anyone, you can try a few prompts in the built-in playground.

Scaling can follow concurrent requests (good for chat and other latency-sensitive traffic), requests per second (good for high-throughput jobs) or custom metrics. If you set active workers to zero, the endpoint scales to zero when there is no traffic, and you pay only while workers are running.

Sizing is the other half of the cost question, and it follows from the weights. A 7B or 13B fine-tune in bf16 is 14-26 GB and fits comfortably on a 48 GB L40S at ₹102/hour. A 70B in bf16 is about 140 GB before the KV cache, so it needs two 80 GB H100s at ₹255.55/hour each — or 4-bit quantisation, which brings it near 40 GB and back onto a single card. Current rates for every GPU are on the [pricing page](https://www.e2enetworks.com/pricing#gpu-pricing).

For private traffic, you can attach the endpoint to a VPC and control access with security groups.

## How to call your endpoint from OpenAI client code

Endpoints on TIR are OpenAI-compatible, so the app team does not need a new SDK. If their code already uses the OpenAI Python client, they change the base URL and the key:

```python
from openai import OpenAI

client = OpenAI(
    base_url="https://<your-endpoint-url>/v1",
    api_key="<your-tir-api-token>",
)

response = client.chat.completions.create(
    model="<your-model-name>",
    messages=[{"role": "user", "content": "Summarise this support ticket in two lines."}],
)

print(response.choices[0].message.content)
```

Streaming is the same call with `stream=True`, and the endpoint also answers plain HTTP calls to `/v1/chat/completions` with a bearer token. WebSocket streaming is supported too. The console shows ready-to-copy cURL and Python samples for each endpoint you create.

This also makes it easy to compare models. If you are currently on a closed model API, point the same code at your own endpoint, run your evaluation set, and compare the results side by side.

## vLLM vs SGLang vs Triton: choosing a serving runtime

For most LLM chat workloads, vLLM is a sensible place to start. The others earn their place in specific situations.

| Runtime | Reach for it when | Good to know |
| --- | --- | --- |
| vLLM | You are serving a popular open LLM for chat or completion | Continuous batching and paged attention keep GPU memory use efficient; supports most Hugging Face LLM architectures |
| SGLang | Many requests share a long prefix, such as a system prompt, agent loops or multi-turn chat | Reuses cached prefixes across requests; strong support for structured (JSON) output |
| Triton | You are serving non-LLM models such as embeddings, classifiers or speech, or mixing frameworks | Runs PyTorch, ONNX and TensorRT models; can chain models into one pipeline |
| Custom container | Your model needs its own pre- or post-processing, pinned dependencies or an unusual architecture | Full control of the stack; you own the image |

TIR also offers PyTorch and NVIDIA Dynamo runtimes. If you are unsure, deploy the same model on two runtimes, send both the same traffic, and compare latency and cost per request. With scale-to-zero, the test endpoint stops billing once its workers scale down after your idle timeout. For a worked example of vLLM on TIR end to end, see [Launching and using Pixtral-12B on TIR](https://www.e2enetworks.com/blog/launching-and-using-pixtral-12b-on-tir-ai-platform-bill-parsing-with-vllm).

## Model inference in India: where your endpoint actually runs

For a lot of Indian teams this is the part that decides the platform, not the runtime.

Endpoints on TIR run in E2E regions in India — Delhi NCR and Chennai — so prompts and responses stay under Indian jurisdiction, and requests from Indian users are not crossing an ocean and back before the first token. Billing is in INR against Indian GPU rates rather than a dollar bill that moves with the exchange rate.

If the traffic should not touch the public internet at all, attach the endpoint to your VPC for a private IP and use security groups to control which ranges can reach it. That covers the common case where prompts carry customer data and a public URL is not acceptable, whatever the auth on it. The platform carries seven security and quality certifications, including SOC 2, ISO 27001 and PCI DSS.

## Get your model off the notebook

A model only starts paying back the time spent training it once other people can use it. If you have one sitting in a notebook right now, the quickest test is to deploy it, send it a few prompts in the playground, and hand the URL to whoever asked for it.

**[Know more about TIR Model Endpoints →](https://www.e2enetworks.com/inference-endpoints)**

Still experimenting before you deploy? [AI Dev Nodes](https://www.e2enetworks.com/ai-dev-nodes) give you a ready GPU workspace with JupyterLab to finish the work first, and [TIR](https://www.e2enetworks.com/tir) covers the rest of the platform — datasets, training jobs and the model catalog. If you are thinking past the first endpoint, [Scaling AI in production](https://www.e2enetworks.com/blog/scaling-ai-in-production) is a good next read.

## Frequently Asked Questions

### How do I deploy a fine-tuned model as an API?

Upload the model to Hugging Face, a GitHub project or your TIR Model Repository, then create a Model Endpoint on TIR. Pick a runtime such as vLLM, choose a GPU, set scaling rules and generate an API token. The endpoint is then callable over HTTPS.

### Should I merge my LoRA adapter before deploying?

If you are serving a single fine-tune, yes — merging the adapter into the base weights gives you one self-contained model that any runtime can serve, with no adapter overhead at inference time. Keep the adapter separate when you are serving several fine-tunes of the same base model and want them behind one deployment instead of one GPU each.

### Can I deploy any Hugging Face model?

You can point an endpoint at a Hugging Face model, as long as the runtime you pick supports its architecture. For models no runtime supports, you can bring your own container.

### Is the endpoint compatible with the OpenAI API?

Yes. Endpoints expose OpenAI-style routes such as /v1/chat/completions with bearer-token authentication, so existing OpenAI client code works after you change the base URL and key.

### What is scale-to-zero, and when should I use it?

Scale-to-zero lets an endpoint shut down all workers when there is no traffic. It suits internal tools, test endpoints and anything with quiet hours. The trade-off is a short warm-up on the first request after an idle period.

### Which is better for serving LLMs, vLLM or SGLang?

Neither wins everywhere. vLLM is a strong general default for chat and completion. SGLang tends to help when many requests share a long prefix, such as agent loops or multi-turn chat. Testing both on your own traffic is the most reliable way to decide.

### Where are the models served from?

Endpoints on TIR run in E2E regions in India (Delhi NCR and Chennai), and you can keep inference traffic on a private network by attaching the endpoint to a VPC.
