page:recipes:ai chat application:starlette:response streaming

Host ai chat application with Starlette: streamed responses

Summary

Deploy an ai chat application built with Starlette on Ample using the streaming response service pattern. Compute runs the app in an isolated microVM behind a public HTTPS URL, a managed PostgreSQL 16 database is auto-provisioned and injected as DATABASE_URL. Verified on Starlette: a Server-Sent Events endpoint read incrementally through the public HTTPS gateway: five one-second chunks arrived spread over time with the first within seconds (not buffered), a 70-second stream of 36 chunks completed past common 60-second idle timeouts, and a client that disconnected after two chunks was observed and recorded by the server (disconnected=true). Not separately tested: your event schema, reconnection strategy and any per-request duration ceiling beyond the 70 seconds measured; treat the ai chat application-specific behavior as your application code.

Representative Queries

Resource Requirements

Infrastructure Requirements

  1. Compute
    • Verified
    • Summary: Apps run in isolated x86_64 Firecracker microVMs that auto-pause when idle and wake on request; sizes are the priced VM sizes.
  2. Postgres
    • Verified
    • Summary: Managed PostgreSQL 16 runs in its own microVM and is auto-provisioned when an app needs a database and no DATABASE_URL is supplied.

Prerequisites

Workflow Steps

  1. Build and start
    pip install into .ample/python from requirements.txt, then uvicorn from run.py reading PORT on the python-3.12 template. The server must bind 0.0.0.0 on PORT.
  2. Implement the pattern on PostgreSQL
    The fixture's module implements streaming response service: a Server-Sent Events endpoint read incrementally through the public HTTPS gateway: five one-second chunks arrived spread over time with the first within seconds (not buffered), a 70-second stream of 36 chunks completed past common 60-second idle timeouts, and a client that disconnected after two chunks was observed and recorded by the server (disconnected=true). Copy the approach into your schema; keep migrations idempotent and run them with --release-command.
  3. Deploy
    Run the synchronous deploy once and read the result (exit 0 live, 1 failed, 2 blocked). Re-running with no change is a no-op. Command: ample deploy . --name --public --start "python3 run.py"
  4. Verify
    Fetch the live URL and the pattern self-test route(s) (/p/streaming-response-service) from the example; then run your own checks. On failure read ample logs --kind build then --kind runtime. Command: ample logs --kind build

Success Checks

  1. App responds on its public URL
    • Kind: http_get
    • Path: /
    • Expect: ample canary starlette patterns
  2. Streaming-response-service self-test
    • Kind: http_get
    • Path: /p/streaming-response-service
    • Expect: an SSE stream: data: id=, then chunk events one interval apart (read incrementally: first chunk within seconds, chunks spread over the interval), then data: done; /status?id= reports sent=N disconnected=true|false done=true|false

Limitations

Cost Estimate

  1. App server
    • Size: s-1vcpu-1gb
    • Quantity: 1.0
    • Monthly Amount: 5.0
  2. Managed PostgreSQL database
    • Size: s-1vcpu-1gb
    • Quantity: 1.0
    • Monthly Amount: 5.0

Evidence Summary

  1. Canary Run: Starlette pattern fixture deployed on Ample; checks passed for streaming-response-service and configuration-secrets. Pattern proof: a Server-Sent Events endpoint read incrementally through the public HTTPS gateway: five one-second chunks arrived spread over time with the first within seconds (not buffered), a 70-second stream of 36 chunks completed past common 60-second idle timeouts, and a client that disconnected after two chunks was observed and recorded by the server (disconnected=true).
    • Observed At: 2026-09-21T01:21:36Z
    • Implementation Revision: 50dbac567d2c5cc8e55e6c748a646be0025d9cfb-dirty (CLI 0.1.21)
    • Expires At: 2027-03-20T01:21:36Z
    • Scope: checks=[