Documentation
Everything needed to call Jaynepal 1.1.
Jaynepal 1.1 is a Nepali-first language model made in Nepal and served by Puja Set Nepal from a hosted, OpenAI-compatible endpoint. This page covers what the model is, how to send it a request, and what it is genuinely not good at yet.
Overview
You reach Jaynepal 1.1 over plain HTTP. There is no SDK to install, no weights to download and no GPU to rent: send a POST to the completions endpoint with your messages, and the answer comes back — streamed by default.
The interface is the OpenAI chat-completions shape, so an existing client library or agent framework works by changing two values: the base URL, and the model id.
Base URL https://pujasetnepal.com/v1 Model id jaynepal-1.1 Endpoint POST https://pujasetnepal.com/v1/chat/completions Models GET https://pujasetnepal.com/v1/models
How a request is served
your app ──► hosted API ──► Jaynepal 1.1 ──► streamed reply
(pujasetnepal.com) (9B)The API is the boundary you integrate against. Which host the model runs on, how it is batched and how it is scaled are ours to change without asking you to redeploy, as long as the contract below holds.
Model card
| Field | Value |
|---|---|
| Model | Jaynepal 1.1 |
| Model id | jaynepal-1.1 |
| Version | 1.1 |
| Parameters | 9B |
| Languages | Nepali (Devanagari) · Romanised Nepali · English |
| Context window | 8,192 tokens |
| Serving interface | OpenAI-compatible HTTP API, SSE streaming |
| Trained by | Samir Puri — Bharatpur, Chitwan, Nepal |
What it is built for
Nepali conversation and Nepali prose. It replies in the language you wrote in — Devanagari, Romanised Nepali or English — and holds that choice instead of drifting back to English. It handles the ordinary work a Nepali-speaking user needs from an assistant: explaining concepts, drafting text, translating between the three registers, and writing code with English technical terms around Nepali explanation.
It also writes Markdown, because structure is what makes a long answer readable — headings, lists, tables and fenced code blocks come back as Markdown rather than as plain paragraphs.
What it is not
Stated plainly, because these are the questions a reader will actually hit:
| Limitation | What that means in practice |
|---|---|
| It is small | 9B parameters, not a frontier model. It is strong for its size in Nepali and weaker than a very large model on hard reasoning, long maths and obscure world knowledge. |
| Academic Nepali is uneven | Formal legal, medical and academic registers are less reliable than everyday conversation. For anything published, have a native speaker review the output. |
| It is not a knowledge base | It can state things confidently and wrongly. Treat factual claims — especially dates, numbers and current events — as drafts to verify. |
| No tool use | It does not browse, call functions or run code. It answers from the prompt and what it learned while training. |
| Devanagari is not transliterated for you | Ask for Romanised Nepali or Devanagari explicitly if the client or channel needs one of them. |
Quickstart
Nothing to install. Export the base URL and your key, then send a request.
export JAYNEPAL_API="https://pujasetnepal.com/v1"
export JAYNEPAL_API_KEY="..." # when access is keyed
curl -N "$JAYNEPAL_API/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $JAYNEPAL_API_KEY" \
-d '{
"model": "jaynepal-1.1",
"messages": [
{ "role": "user", "content": "लमजुङ किन प्रसिद्ध छ?" }
],
"stream": true
}'Confirm the model is there
curl "$JAYNEPAL_API/models"
{ "object": "list",
"data": [ { "id": "jaynepal-1.1", "object": "model" } ] }From Python, with an OpenAI client
from openai import OpenAI
client = OpenAI(base_url=os.environ["JAYNEPAL_API_URL"],
api_key=os.environ["JAYNEPAL_API_KEY"])
stream = client.chat.completions.create(
model="jaynepal-1.1",
messages=[{"role": "user", "content": "नमस्ते, तपाईं कस्तो छ?"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)API reference
POST /chat/completions
POST https://pujasetnepal.com/v1/chat/completions
Content-Type: application/json
Authorization: Bearer <key> # when access is keyed
{
"model": "jaynepal-1.1",
"messages": [
{ "role": "system", "content": "optional persona" },
{ "role": "user", "content": "नमस्ते" }
],
"temperature": 0.7,
"max_tokens": 1024,
"stream": true
}Request body
| Field | Type | Default | Notes |
|---|---|---|---|
model | string | jaynepal-1.1 | Required. The only id currently served. |
messages | array | — | Required. Roles are user, assistant and system. Send the whole conversation each time; the model holds no state between requests. |
temperature | number | 0.7 | 0 is nearly deterministic, 1 is loose. Lower it for factual or technical answers, raise it for drafting. |
max_tokens | integer | 1024 | Ceiling on the reply. Generous values cost latency you do not get back if the model finishes early. |
stream | boolean | true | Server-sent events instead of one JSON body. See Streaming below. |
Response — non-streaming
{
"id": "chatcmpl-...",
"object": "chat.completion",
"model": "jaynepal-1.1",
"choices": [
{
"index": 0,
"finish_reason": "stop",
"message": { "role": "assistant", "content": "..." }
}
],
"usage": { "prompt_tokens": 18, "completion_tokens": 96, "total_tokens": 114 }
}GET /models
Lists what the endpoint is serving. Useful as a cheap liveness check, and as the first thing to try when a request fails.
GET https://pujasetnepal.com/v1/models
{ "object": "list", "data": [ { "id": "jaynepal-1.1", "object": "model" } ] }Streaming
With stream: true — the default — the response is a sequence of server-sent events, each carrying one fragment of the reply, terminated by a literal [DONE] line.
data: {"choices":[{"delta":{"content":"लमजुङ"}}]}
data: {"choices":[{"delta":{"content":" जिल्ला"}}]}
data: {"choices":[{"delta":{"content":" हो"}}]}
data: [DONE]Consuming it without breaking Nepali
Read the response as a byte stream and decode with buffering enabled, splitting lines only on complete newlines. A Devanagari character is several bytes, and a token can land in the middle of one; decoding each chunk independently turns that into mojibake in exactly the language this model exists to serve.
reader = response.iter_lines(decode_unicode=False)
buffer = b""
for chunk in reader:
buffer += chunk + b"\n"
lines = buffer.split(b"\n")
buffer = lines.pop() # keep the partial line for next time
for line in lines:
if not line.startswith(b"data: "):
continue
payload = line[6:]
if payload == b"[DONE]":
break
text = json.loads(payload)["choices"][0]["delta"].get("content", "")
print(text, end="", flush=True)Errors
Failures are JSON, not an empty stream — so a client can tell “the model said nothing” apart from “the model could not be reached”.
{
"error": {
"message": "human-readable description",
"type": "invalid_request_error",
"code": "bad_request"
}
}| Status | Code | Cause |
|---|---|---|
| 400 | bad_request | Body was not JSON, messages was empty, or a message had no valid role and string content. |
| 401 | invalid_api_key | Missing or wrong key, when the endpoint is keyed. |
| 404 | model_not_found | The model field does not name the served model. |
| 429 | rate_limited | Too many requests in flight. Back off and retry. |
| 502 | upstream_unreachable | The model host could not be reached — usually the serving session not running. |
| 503 | unavailable | The model is loading. Retry shortly. |
| 504 | timeout | No response within the gateway window. Long answers on a CPU-only host can hit this. |
429 and 503 with exponential backoff. Do not retry 400 — the request itself is wrong and will fail again.Access and status
The canonical production base URL is https://pujasetnepal.com/v1. Access is arranged directly at the moment rather than granted by a signup form, and the endpoint you are given is the one to use in base_url.
Bringing up an endpoint yourself
The interface is standard, so any server that speaks OpenAI chat completions can front the model while the hosted one is being brought up — a rented GPU box with vLLM, llama.cpp, or a GPU Space. Point your client at it and the code below does not change.
# 1. serve it vllm serve <jaynepal-1.1> --port 8000 --max-model-len 8192 # 2. confirm the contract curl http://localhost:8000/v1/models # 3. use it — the only change your client needs export JAYNEPAL_API=http://localhost:8000/v1
Limits and roadmap
The gaps, stated as gaps. A reader deciding whether to build on a model needs these more than it needs a feature list.
| Today | Consequence |
|---|---|
| No self-serve keys | Access is arranged directly, so onboarding is not instant. |
| The evaluation endpoint is session-bound | It is offline between GPU sessions, and its address moves. |
| No published benchmark | Quality claims are qualitative until the evaluation below ships. |
| 8,192 tokens context | Very long documents must be chunked or summarised rather than pasted whole. |
| No tool calling | It cannot browse, run code or call your functions. |
Roadmap
- A permanently hosted endpoint on a persistent GPU host.
- Self-serve API keys with per-key quotas and usage reporting.
- A published Nepali evaluation — comprehension, generation and Romanised transcription — against comparable models.
- A quantised release for people who want to run it on their own hardware.
- Deeper register coverage for formal, legal and academic Nepali.
Credits
Jaynepal 1.1 and this platform are built by Samir Puri (DevSamirX) in Bharatpur, Chitwan, Nepal.
The model is published on Hugging Face. The source for this website is on GitHub under the MIT licence, and issues and pull requests are welcome.