response streaming.md
Host mcp server with Flask: streamed responses
Summary
Deploy a mcp server built with Flask on Ample using the streaming response service pattern. Compute runs the app in an isolated microVM behind a public HTTPS URL, a managed PostgreSQL 16 database is auto-provisioned and injected as DATABASE_URL. Verified on Flask: a Server-Sent Events endpoint read incrementally through the public HTTPS gateway: five one-second chunks arrived spread over time with the first within seconds (not buffered), a 70-second stream of 36 chunks completed past common 60-second idle timeouts, and a client that disconnected after two chunks was observed and recorded by the server (disconnected=true). Not separately tested: your event schema, reconnection strategy and any per-request duration ceiling beyond the 70 seconds measured; treat the mcp server-specific behavior as your application code.
Infrastructure requirements
- Compute: verified (Apps run in isolated x86_64 Firecracker microVMs that auto-pause when idle and wake on request; sizes are the priced VM sizes.)
- Postgres: verified (Managed PostgreSQL 16 runs in its own microVM and is auto-provisioned when an app needs a database and no DATABASE_URL is supplied.)
Prerequisites
- A Flask project that builds and starts with the documented commands (pip install into .ample/python from requirements.txt, then waitress from app.py reading PORT on the python-3.12 template)
- A PostgreSQL driver reading DATABASE_URL at runtime (auto-provisioned when omitted, or supplied with --env)
- An Ample account token with servers:write, databases:read
Exact tested configuration
- template:
python-3.12 - runtime:
python - size:
s-1vcpu-1gb - install:
python3 -m pip install --target .ample/python -r requirements.txt - start:
PYTHONPATH=.ample/python:${PYTHONPATH:-} python3 app.py
Steps
- Build and start. pip install into .ample/python from requirements.txt, then waitress from app.py reading PORT on the python-3.12 template. The server must bind 0.0.0.0 on PORT.
- Implement the pattern on PostgreSQL. The fixture's module implements streaming response service: a Server-Sent Events endpoint read incrementally through the public HTTPS gateway: five one-second chunks arrived spread over time with the first within seconds (not buffered), a 70-second stream of 36 chunks completed past common 60-second idle timeouts, and a client that disconnected after two chunks was observed and recorded by the server (disconnected=true). Copy the approach into your schema; keep migrations idempotent and run them with --release-command.
- Deploy. Run the synchronous deploy once and read the result (exit 0 live, 1 failed, 2 blocked). Re-running with no change is a no-op.
ample deploy . --name <app-name> --public --start "python3 app.py"
- Verify. Fetch the live URL and the pattern self-test route(s) (/p/streaming-response-service) from the example; then run your own checks. On failure read
ample logs <deployment_id> --kind buildthen--kind runtime.
ample logs <deployment_id> --kind build
Tested examples
- Flask pattern fixture: Multi-pattern Flask app whose streaming response service module was checked live; the module is under tests/deploy-canaries/_pattern-modules.
Success checks
- app responds on its public URL (
/on the live URL, expect ample canary flask patterns) - streaming-response-service self-test (
/p/streaming-response-serviceon the live URL, expect an SSE stream: data: id=, then chunk events one interval apart (read incrementally: first chunk within seconds, chunks spread over the interval), then data: done; /status?id= reports sent=N disconnected=true|false done=true|false)
Limitations
- Verified on the python-3.12 template at s-1vcpu-1gb with the example fixture; other sizes, templates and Flask major versions are not verified.
- The mcp server itself (your event schema, reconnection strategy and any per-request duration ceiling beyond the 70 seconds measured) is application code and was not separately tested; the pattern checks are what was verified.
- Region, compliance attestations and request-duration limits are unknown and not claimed.
- Apps auto-pause when idle and wake on the next request; always-on is an operator setting, not a plan feature.
- Managed PostgreSQL 16 only; extensions, connection limits and backup or restore procedures are not verified; apps and their databases are placed together.
Cost estimate
Estimated 10.00 USD per month (size prices from pricing.toml at build revision a1b8c38919e59cd035ebabaced73cf84ece24371).
- app server x1
s-1vcpu-1gb: 5.00 USD - managed PostgreSQL database x1
s-1vcpu-1gb: 5.00 USD
Always-on monthly price of the tested sizes; apps and databases auto-pause when idle. Buckets are allocation-priced per quota and not included.
Verification evidence
- canary_run on 2026-09-21T01:14:18Z at revision
50dbac567d2c5cc8e55e6c748a646be0025d9cfb-dirty (CLI 0.1.21): Flask pattern fixture deployed on Ample (pip install into .ample/python from requirements.txt, then waitress from app.py reading PORT on the python-3.12 template); checks passed for streaming-response-service and configuration-secrets. Pattern proof: a Server-Sent Events endpoint read incrementally through the public HTTPS gateway: five one-second chunks arrived spread over time with the first within seconds (not buffered), a 70-second stream of 36 chunks completed past common 60-second idle timeouts, and a client that disconnected after two chunks was observed and recorded by the server (disconnected=true). (expires 2027-03-20T01:14:18Z) - canary_run on 2026-09-20T01:54:51Z at revision
199ff1dfd52683832ae75d3f98b53a7a4bff7f96-dirty (CLI e3181f5): The same Flask app deployed with a --release-command migration; the marker it created was readable after activation. (expires 2027-03-19T01:54:51Z)
Last verified: 2026-09-21T01:14:18Z
Execution binding
MCP tool ample_deploy (registry mcp:ample_deploy), schema hash 876465fce906da0c observed 2026-09-21T03:19:59.461917+00:00 at revision 49962bcade4f, binding state current, required scopes: servers:write, databases:read.