Event ingestion
POST /events and the queue behind it — every SDK write and every HTTP batch lands here.
// status
This is a published page, not a live health check. It shows what we knew when a person last updated it on 5 September 2026 — no monitor writes to it automatically. If something looks wrong right now, email support and we will tell you what we are seeing.
Each surface is tracked on its own, because they fail on their own. Uptime is calculated from the incidents recorded further down this page, not from an external monitor.
POST /events and the queue behind it — every SDK write and every HTTP batch lands here.
Authenticated reads on api.alitycs.com/v1 — events, funnels, retention, and exports.
Ask Data: plain-language questions, the query planner behind them, and the answers it returns.
app.alitycs.com — sign-in, charts, saved reports, and workspace settings.
Outbound delivery of workspace events to your endpoints, including the retry schedule.
Thirty-day figures cover 7 Aug 2026 to 5 Sep 2026. Both figures count degraded time as downtime, so a slow component reads the same as an unavailable one.
| Component | Current state | 30-day uptime | 90-day uptime | Incidents in window |
|---|---|---|---|---|
| Event ingestion | Operational | 100.00% | 99.98% | 1 resolved |
| Query API | Operational | 99.89% | 99.96% | 1 resolved |
| AI agent | Operational | 100.00% | 99.95% | 1 resolved |
| Dashboard | Operational | 99.89% | 99.96% | 1 resolved |
| Webhooks | Operational | 100.00% | 99.85% | 1 resolved |
Everything that affected customer traffic in this window, with what caused it and what changed afterwards. Every one of them is resolved.
19 August 2026 · 14:12–14:59 UTC
A materialized view for daily event rollups was deployed before its backfill had finished. Queries covering the current day fell through to a full scan of the raw event table, so p95 read latency rose from 380 ms to just over 9 s and roughly 6% of requests hit the 30-second gateway timeout. Dashboard charts that depend on those reads showed a load error rather than stale numbers. Writes were unaffected and no events were lost.
Resolution. Query planning was routed back to the previous view and latency recovered within four minutes. Deploys now gate view routing on a backfill-completion check, so a partially built view cannot take read traffic.
28 July 2026 · 09:41–10:03 UTC
One workspace replayed fourteen days of history as unbatched single-event requests, about forty times its normal write rate. The ingest limiter shed load across the shared queue instead of against that workspace alone, so unrelated workspaces saw 429 responses for 22 minutes. SDK retry and backoff absorbed most of it; clients that had disabled retries dropped events inside the window.
Resolution. The replaying workspace was throttled and error rates returned to baseline. Ingest quotas are now enforced per workspace with a ceiling of 20% of shared capacity, so one backfill can no longer displace another workspace's traffic.
6 July 2026 · 18:20–19:28 UTC
A query-planner change sent the complete event schema to the model instead of a relevant subset. Workspaces with more than roughly 400 distinct event names exceeded the model's context limit, and every question they asked returned an error; smaller workspaces were unaffected. Saved reports and the REST API kept working, since neither goes through the planner.
Resolution. The planner change was reverted. The schema handed to the model is now selected by relevance and truncated deterministically, and a contract test exercises a workspace with 1,000 event names on every planner change.
15 June 2026 · 03:05–06:17 UTC
A customer endpoint started answering after 25 seconds rather than failing outright. The delivery worker treated a slow success as a retry candidate and re-queued it, and because deliveries drained through one shared queue in order, unrelated endpoints waited behind the backlog. Nothing was dropped — the queue drained completely once the backlog cleared.
Resolution. Delivery timeouts are now 10 seconds, a slow success is recorded as a success, and every endpoint has its own queue with exponential backoff. A single slow consumer now delays only its own deliveries.
We do not publish a live maintenance calendar here, because there is nothing feeding one. This is the policy the team works to instead.
Planned work happens Tuesdays between 02:00 and 04:00 UTC, the quietest two hours across the regions we serve. Most releases need no window at all and ship without one.
Anything expected to interrupt reads or writes is emailed to workspace owners at least seven days ahead, with the window, the components involved, and what to expect.
Writes are queued rather than rejected during maintenance and drain when it ends, so you do not lose events. Reads and the dashboard may be briefly unavailable.
Security fixes can move without seven days of notice. When that happens we email first and write it up here afterwards, in the incident list above.
Most of what reaches us turns out to be a key, a date range, or a blocked request. The troubleshooting guide covers those in order.
Still stuck? support@alitycs.com — replies within one business day.