Version 1.18.1.4
Release Date: September 4, 2026
Release Type: Stable
Previously published as: 1.18.9 (image tag staging-1.18.9)
Store outage history with a guided resync, a complete record of every log deletion run, optional limits on the PostgreSQL monitoring fallback, and Health Check alerting and retention fixes.
Backend Server
New Features
- Nightly log deletion records its outcome: The scheduled log deletion now writes a full record of each run: the period it covered, when it started and finished, and how many PostgreSQL rows and Elasticsearch documents it removed. A run ends as COMPLETED with what it removed, SKIPPED with the reason it had nothing to do, or FAILED with the actual error message. When the deletion could not be confirmed, the figures are reported as a lower bound instead of a zero.
- Optional limits on the PostgreSQL monitoring fallback: While Elasticsearch is unavailable, monitoring records are written to PostgreSQL so that message processing continues. Two optional limits can now bound this fallback: a maximum total size of the fallback tables, including their indexes, and a maximum outage duration. When a limit is reached, further monitoring records are dropped while messages keep being processed; this is logged once, and recording resumes on its own as soon as the nightly deletion frees space or Elasticsearch returns. Both limits are unset by default, so existing installations behave exactly as before.
Bug Fixes
- A failed nightly deletion keeps its record: When the scheduled log deletion failed, its record was removed, so the failure left no trace. The record is now kept and marked FAILED.
- Deletion period end date names the last day deleted: The end date recorded for a deletion run is now the last day that was deleted rather than the day after it, consistent with the other platform services.
Frontend Server
New Features
- Log deletion jobs report results per table: A deletion job now reports how many records it removed from each table and index, in PostgreSQL and in Elasticsearch. A table without a figure was not measured on that run, which is different from a table measured at zero. A cancelled job records where it stopped, and a failed job carries a summary of the failure.
- In-flight message details are part of store synchronization: The in-flight message detail records (delivering log details) are now a sync target and can be transferred in both directions between PostgreSQL and Elasticsearch, so a store outage no longer leaves a gap in them that cannot be repaired.
Bug Fixes
- Automatic resync after an Elasticsearch outage now runs: When Elasticsearch became available again, the automatic resync job was created but never started and stayed PENDING indefinitely; only manually started sync jobs ran. The automatic job now starts and runs to completion.
- Deletion jobs no longer over-report or finish as completed after an error: For a period containing both expired and still-retained days, a deletion job counted the rows of the retained days as deleted although they were still present. Counts now cover only the days that were actually deleted, and a job in which a day failed can no longer end as COMPLETED.
Frontend Web
New Features
- Store Outages table with guided resync: The Sync Jobs page gains a Store Outages table with one row per store instance, so each Elasticsearch node and each PostgreSQL replica is listed separately. Its action opens the Start Sync dialog pre-filled with the outage period to the minute and with the direction pointing into the store that was down. A banner on Database Management, System Health and Sync Jobs points out an ended Elasticsearch outage that has not been resynced yet.
- Sync direction named in words: The Start Sync dialog now offers the direction as a selection that names both stores in words instead of arrow buttons.
- Data store instances on System Health: The live per-instance table moved from Database Management to System Health, next to the component cards. The cards show whether a component is up; the table shows which node, pod or instance is affected.
- Job lists with one progress and status column: Deletion Jobs, Sync Jobs and AI Monitoring show progress and status in a single column: a running job shows how far it has got, a finished job shows how it ended. A running job whose progress is unknown shows a dash instead of a full bar labelled 0%.
- Detail views for deletion and sync jobs: Deletion jobs and sync jobs each have a detail view that shows the outcome per table. A figure that was not measured is shown as not measured rather than as 0.
- Sorting and search on Sync Jobs, paging on Flow Retention: The Sync Jobs list can now be sorted by its columns and searched by job ID and error message, and the Flow Retention list is divided into pages.
Bug Fixes
- All pages of Sync Jobs and Log Cleaning are reachable: The pager on these lists did not receive the number of pages, so only the first page of jobs could be opened although the total was displayed correctly. The pager now works.
- System Logs page is displayed again: The System Logs page rendered empty. The page content is shown again.
Health Check
New Features
- Store outage history: Health Check now records an outage for each store instance. An outage is opened when the failure threshold is reached and closed with the time of the first healthy check; outages left open by a service restart are reconciled at startup. For each outage a suggested sync period is calculated that points into the store that was down, which feeds the new Store Outages table.
- Log deletion by Health Check is recorded: Every log deletion that Health Check performs now writes a full execution record. If PostgreSQL is unreachable at that moment, the record is kept on disk, so an outage does not cost the audit trail and recording never blocks the deletion itself.
- Separate threshold for database server connection saturation: The alert for database server connection saturation now has its own threshold instead of sharing the connection pool threshold. The existing value is carried over automatically at startup.
Bug Fixes
- Healthy Elasticsearch 7.17 clusters are no longer reported as unreachable: A statistics call that could not be read from Elasticsearch 7.17 caused the whole check to fail, so the component was shown as DOWN while the cluster was answering, and the cluster status, unassigned shard and store size values that alerts depend on were never published. Only the health call is required now; the additional statistics are collected separately and stay empty when they are unavailable.
- Log retention sweep runs every day: The retention sweep ran only once after the service started and never again for as long as the service kept running, so logs stopped ageing out. It now runs on every scheduled day.
- Alert mails are sent while PostgreSQL is down: Alert recipients were read from the database at send time, so when PostgreSQL itself was unavailable no alert mail went out at all. Recipients now fall back to the last successfully read list, and to the configured
alerts.email.toaddress when no list has been read yet. An intentionally empty recipient list stays empty. - Deletion period end date names the last day deleted: The end date recorded for a deletion run is now the last day that was deleted, consistent with the other platform services.
1.18.1.4 is a stable release. Previous: 1.18.1.3