page:recipes:ai chat application:spring boot:response streaming
Host ai chat application with Spring Boot: streamed responses
Deploy an ai chat application built with Spring Boot on Ample using the streaming response service pattern. Compute runs the app in an isolated microVM behind a public HTTPS URL, a managed PostgreSQL 16 database is auto-provisioned and injected as DATABASE_URL. Verified on Spring Boot: a Server-Sent Events endpoint read incrementally through the public HTTPS gateway: five one-second chunks arrived spread over time with the first within seconds (not buffered), a 70-second stream of 36 chunks completed past common 60-second idle timeouts, and a client that disconnected after two chunks was observed and recorded by the server (disconnected=true). Not separately tested: your event schema, reconnection strategy and any per-request duration ceiling beyond the 70 seconds measured; treat the ai chat application-specific behavior as your application code.
Representative Queries
- Host ai chat application with Spring Boot: streamed responses
- Where can I host AI chat application built with Spring Boot?
- I need a tested end-to-end streaming transport, timeout behavior and client disconnect handling.
Resource Requirements
- Compute: Apps run in isolated x86_64 Firecracker microVMs that auto-pause when idle and wake on request; sizes are the priced VM sizes.
- Postgres: Managed PostgreSQL 16 runs in its own microVM and is auto-provisioned when an app needs a database and no DATABASE_URL is supplied.
Framework
Spring Boot
Workload
AI chat application
Release Status
Published
Support Status
Verified
Execution Status
Ready
Prerequisites
- A Spring Boot project that builds and starts with the documented commands (./mvnw -q -DskipTests package (or the Gradle wrapper) producing one application jar, then java -jar on the jvm-21 template (Temurin JDK 21) with server.port read from PORT)
- A PostgreSQL driver reading DATABASE_URL at runtime (auto-provisioned when omitted, or supplied with --env)
- An Ample account token with servers:write, databases:read
Tested Configuration
- Template: jvm-21
- Runtime: java
- Size: s-1vcpu-2gb
- Build: ./mvnw -q -DskipTests package
- Start: java -jar $(ls target/*.jar | grep -v -- '-plain' | head -n 1)
Workflow Steps
- Build and start
./mvnw -q -DskipTests package (or the Gradle wrapper) producing one application jar, then java -jar on the jvm-21 template (Temurin JDK 21) with server.port read from PORT. The server must bind 0.0.0.0 on PORT. - Implement the pattern on PostgreSQL
The fixture's module implements streaming response service: a Server-Sent Events endpoint read incrementally through the public HTTPS gateway: five one-second chunks arrived spread over time with the first within seconds (not buffered), a 70-second stream of 36 chunks completed past common 60-second idle timeouts, and a client that disconnected after two chunks was observed and recorded by the server (disconnected=true). Copy the approach into your schema; keep migrations idempotent and run them with --release-command. - Deploy
Run the synchronous deploy once and read the result (exit 0 live, 1 failed, 2 blocked). Re-running with no change is a no-op.ample deploy . --name <app_name> --public - Verify
Fetch the live URL and the pattern self-test route(s) (/p/streaming-response-service) from the example; then run your own checks. On failure readample logs --kind buildthen--kind runtime.ample logs --kind build
Success Checks
- App responds on its public URL
http_get /
Expect: ample canary spring boot patterns - Streaming-response-service self-test
http_get /p/streaming-response-service
Expect: an SSE stream: data: id=, then chunk events one interval apart (read incrementally: first chunk within seconds, chunks spread over the interval), then data: done; /status?id= reports sent=N disconnected=true|false done=true|false.
Limitations
- Verified on the jvm-21 template at s-1vcpu-2gb with the example fixture; other sizes, templates and Spring Boot major versions are not verified.
- The ai chat application itself (your event schema, reconnection strategy and any per-request duration ceiling beyond the 70 seconds measured) is application code and was not separately tested; the pattern checks are what was verified.
- Region, compliance attestations and request-duration limits are unknown and not claimed.
- Apps auto-pause when idle and wake on the next request; always-on is an operator setting, not a plan feature.
- Managed PostgreSQL 16 only; extensions, connection limits and backup or restore procedures are not verified; apps and their databases are placed together.