Building apps

AI in an app

The backend LLM, tool loops, streaming to the browser, and when to hand off to a managed agent instead.

There is no in-page LLM. llm runs in the backend function, so the browser never holds a model call.

Backend LLM or managed agent?

The line is drawn by what the feature has to do:

The AI feature…Use
Summarizes or classifies data the backend already reads, in secondsBackend llm
Reads, understands, extracts, transforms, or generates any fileManaged agent
Writes and runs codeManaged agent (sandbox)
Must survive the request, be retried, or take minutesManaged agent
Is triggered from Slack or by a scheduleManaged agent
Needs a run history someone will auditManaged agent

The two compose: the backend function runs the fast turns itself and hands heavy steps to agents.start().

Calling the model

const r = await llm.generate({ messages: [{ role: "user", content: "..." }] });
const stream = await llm.stream({ messages });     // for await (const ev of stream)
await llmProviders();                              // what the org has configured

Text in, text out. There are no embeddings, vector search, or multimodal input.

Manifest: llm: true. There is a per-app daily token cap. Exceeding it returns a typed 429.

An org can configure many models across many providers. Calls route by (provider, model):

  • Pass opts.model to pick a catalog model by name. Its provider is implied.
  • Pass opts.provider alone to use that provider's default model.
  • Pass neither to use the org default.
railcode llm providers      # configured providers, each with its models
railcode llm models         # flat list of every callable model

Tool loops

Tools that carry a run make the SDK drive the loop. It validates each call's args against the schema, executes run in your backend function, feeds summarize(result) back, and repeats until the model answers.

await llm.generate({ messages, tools });                    // resolves with the finished answer
for await (const ev of llm.stream({ messages, tools })) {}  // text + step events, live

llm.streamRaw() is the deliberate exception. It hands you the ndjson bytes to relay straight to a browser, so nobody is left in the backend function to execute a run. It refuses tools that carry a run and accepts definitions without one. That is exactly the relay a browser-side loop needs.

Requires @railcode/sdk 0.3.0 or later. Earlier builds routed the internal stream through streamRaw, so every streamed tool loop died on its first turn with tool_loop_error.

Streaming to the frontend

Never hand-roll the ReadableStream. toNdjson(source, opts?) turns any iterable of JSON values into an ndjson Response that you return straight from a route:

app.post("/api/chat", async (c) => {
  const { messages } = await c.req.json();
  return toNdjson(llm.stream(messages, { tools }));
});

It handles the two things that are easy to get wrong:

  • A mid-stream failure cannot be an HTTP status, because the 200 is already sent. It becomes a terminal {"type":"error", error, message} frame. errorFrame() keeps the platform's typed code (daily_token_limit_exceeded, provider_auth_error, and so on), so the browser maps it to advice exactly as on a non-streamed call.
  • A client that hangs up must stop the work. Cancelling closes the source generator, so an abandoned run stops spending tokens.

It takes any iterable. So a route that mixes in its own frames and persists the turn is still one call: write an async generator and return toNdjson(frames()).

The browser side is about 20 lines: read, split on \n, JSON.parse each line, and keep a buffer because a network chunk can split a line. Copy it from apps/chat/frontend/src/lib/api.ts in the examples repo.

Delegating to a managed agent

The backend function handles fast turns and hands off anything that must outlive the request.

app.post("/api/extract", async (c) => {
  const { file } = await c.req.json();
  const run = await agents.start("extractor", { file });     // returns QUEUED immediately
  await db.collection("jobs").put(run.request_id, {
    owner: ctx.user!.id, file, status: "running",
  });
  return c.json({ requestId: run.request_id }, 202);
});

app.get("/api/extract/:id", async (c) => {
  const job = await db.collection("jobs").get(c.req.param("id"));
  if (!job || job.owner !== ctx.user!.id) return c.json({ error: "not found" }, 404);
  const run = await agents.get(c.req.param("id"));
  return c.json({ status: run.status, output: run.output_json });
});

The frontend polls /api/extract/:id.

Runs are never awaited while the request is open. A backend invocation is one request with a limited subrequest budget and a token that expires. An agent run is minutes of work. There is no backend call that waits, and a get() loop is not a substitute: each poll spends a subrequest, and the token expires before a long run finishes.

Rules to remember:

  • Declare the agent. agents: [name] in manifest.yaml. A missing declaration is a 403 refusal rather than pass-through, even for an agent the caller could invoke from the dashboard. An agent that does not exist is a 404.
  • A run is owned by (app, caller). agents.get() reads only runs that this app started for this caller. A run started from the dashboard is invisible to the backend function.
  • Cron cannot start or poll a run. No caller means no run owner (409 both ways). Give the agent its own schedule instead.
  • Prefer org agents. An org agent's app_data_write lands in the app's shared scope, which is your flat store. Results appear in a collection your backend function already reads, with no bridge.

Email

await email.send({ to, subject, html });

Send-only, with a platform-pinned sender and an appended disclaimer. You cannot receive email or send from a custom address. When mail must come from a specific person's own account, use a Gmail connector they own.

Manifest: email: true. Per-app daily cap; exceeding it returns a typed 429.

On this page