fix(o11y): key the deep scrub verdict on lateness, not interval age (#645)

* fix(o11y): key the deep scrub verdict on lateness rather than interval age

* fix(o11y): base the scrub verdict on ceph's latest target

* fix(o11y): note that monk withholds late and breach until a read
This commit is contained in:
Andy Molenda
2026-09-11 09:23:42 -07:00
committed by GitHub
parent 1b00d89432
commit 85273ac719
2 changed files with 483 additions and 76 deletions
+19 -7
View File
@@ -154,8 +154,10 @@ spec:
- evaluator:
params: [0]
type: gt
# Stale thresholds judge overdue silently: the exporter keeps last-known
# intervals when the cluster read fails, which is correct and invisible.
# Stale thresholds judge silently: a failed read keeps the last-known osd
# and mgr config. Before an instance's first successful read monk has only
# ceph's built-in defaults, so it withholds late and breach and judges
# overdue alone against them.
- uid: monk-interval-read-failing
title: MonkIntervalReadFailing
condition: B
@@ -169,8 +171,14 @@ spec:
summary: monk cannot read the cluster's scrub intervals
description: >-
ceph_scrub_interval_read_success is 0 on {{ $labels.instance }}
(cluster {{ $labels.cluster }}) for an hour; overdue thresholds
no longer track the cluster's scrub configuration.
(cluster {{ $labels.cluster }}) for an hour. After a successful read
monk keeps the last-known osd and mgr scrub config, which goes stale
if the cluster's config changes. If this instance has not read the
cluster since it started, it withholds late and breach and judges
overdue against ceph's built-in defaults (7d for both intervals),
which overstates deep overdue on a cluster tuned to 28d; the verdict
can then read Work outstanding, but not Falling behind or Past warn
deadline, from this instance.
data:
- refId: A
relativeTimeRange:
@@ -199,8 +207,10 @@ spec:
- evaluator:
params: [0]
type: gt
# A stamp-format drift below the outage threshold: some PGs unparsable
# (counted overdue, fail-safe) while the refresh still succeeds.
# Unparsable stamps count toward overdue but toward neither
# ceph_scrub_late_* nor ceph_scrub_breach_*, so both undercount while
# this fires; treat the Falling behind and Past warn deadline verdicts
# as unreliable until fixed.
- uid: monk-parse-errors
title: MonkParseErrors
condition: B
@@ -215,7 +225,9 @@ spec:
description: >-
ceph_scrub_parse_errors is nonzero on {{ $labels.instance }}
(cluster {{ $labels.cluster }}): some PG stamps no longer parse
(format drift?); affected PGs count as overdue until fixed.
(format drift?); affected PGs count as overdue but toward neither
ceph_scrub_late_* nor ceph_scrub_breach_*, so treat the Falling
behind and Past warn deadline verdicts as unreliable until fixed.
data:
- refId: A
relativeTimeRange:
+462 -67
View File
@@ -41,13 +41,13 @@
"title": "What this covers",
"gridPos": {
"h": 7,
"w": 19,
"w": 14,
"x": 0,
"y": 0
},
"options": {
"mode": "markdown",
"content": "Ceph guards against silent data corruption by **scrubbing**: re-reading stored data and checking it is still intact. A **regular** scrub compares the bookkeeping of every copy; a **deep** scrub reads all the data back and verifies checksums, which is the slow, disk-heavy part. Every slice of data (a **PG**) must get a regular scrub every 7 days and a deep scrub every 28 days on spice. At petabyte scale a deep cycle takes weeks, so this board answers three questions: what is scrubbing right now, which hosts are busy with it, and is the cycle keeping up.\n\n**Now / PGs**: how many PGs the cluster flags as scrubbing. A PG is flagged from the moment it is queued, so these counts run ahead of the scrubs actually reading data.\n\n**Hosts**: a host counts as scrubbing when one of its disks is leading a scrub that made progress in the last 5 minutes. The data being checked is spread across many hosts, so other hosts read for it too; the read-throughput panel shows that load.\n\n**Work outstanding**: measured, not estimated. monk reads every PG's last-scrub timestamp on the mon hosts and reports, per pool, how much data is older than the target interval it reads from the cluster's own config. 100% coverage means nothing is overdue. Byte figures in this row are stored (logical) bytes and must not be compared with raw capacity without the per-pool expansion factor.\n\n**What good looks like**: the status block reads *On track*, cycle coverage is 100.00%, and no PGs are overdue. Coverage is weighted by stored bytes, which exclude omap, so the PGs overdue count is the one that catches an index pool falling behind. Orange or yellow bands across the charts mark periods when scrubbing was deliberately paused, which is the usual explanation for a flat line."
"content": "Ceph guards against silent data corruption by **scrubbing**: re-reading stored data and checking it is still intact. A **regular** scrub compares the bookkeeping of every copy; a **deep** scrub reads all the data back and verifies checksums, which is the slow, disk-heavy part. Every slice of data (a **PG**) must get a regular scrub every 7 days and a deep scrub every 28 days on spice. At petabyte scale a deep cycle takes weeks, so this board answers three questions: what is scrubbing right now, which hosts are busy with it, and is the cycle keeping up.\n\n**Now / PGs**: how many PGs the cluster flags as scrubbing. A PG is flagged from the moment it is queued, so these counts run ahead of the scrubs actually reading data.\n\n**Hosts**: a host counts as scrubbing when one of its disks is leading a scrub that made progress in the last 5 minutes. The data being checked is spread across many hosts, so other hosts read for it too; the read-throughput panel shows that load.\n\n**Work outstanding**: measured, not estimated. monk reads every PG's last-scrub timestamp on the mon hosts and reports, per pool, how much data is older than the target interval it reads from the cluster's own config. 100% coverage means nothing is overdue. Byte figures in this row are stored (logical) bytes and must not be compared with raw capacity without the per-pool expansion factor.\n\n**What good looks like**: the status block reads *On track* or *Work outstanding*, Past ceph's latest target reads *None* or stays green (under a day past it is normal queue wait), Past warn deadline reads *None*, and deep coverage sits below 100% by the overdue floor. Coverage is weighted by stored bytes, which exclude omap, so an omap-only index pool can fall behind without moving it; the PG counts catch that, with PGs overdue as work outstanding and Past warn deadline as the verdict. Orange or yellow bands across the charts mark periods when scrubbing was deliberately paused, which is the usual explanation for a flat line."
}
},
{
@@ -59,9 +59,9 @@
"uid": "$datasource"
},
"gridPos": {
"h": 7,
"w": 5,
"x": 19,
"h": 4,
"w": 10,
"x": 14,
"y": 0
},
"targets": [
@@ -74,7 +74,7 @@
"instant": true,
"range": false,
"legendFormat": "",
"expr": "sum(max by (cluster, pool_id) (ceph_scrub_overdue_pgs{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"})) / sum(max by (cluster, pool_id) (ceph_scrub_pool_pgs{project=\"yucca\", cluster=~\"$cluster\"}))"
"expr": "(vector(4) and (sum(max by (cluster, pool_id) (ceph_scrub_breach_pgs{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"})) > 0))\nor (3 * (max(ceph_osd_flag_nodeep_scrub{project=\"yucca\", cluster=~\"$cluster\"}) == 1) and on () count(ceph_scrub_overdue_pgs{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"}))\nor (vector(2) and (max(max by (cluster, pool_id) (ceph_scrub_late_max_seconds{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"})) > 86400))\nor (vector(1) and (sum(max by (cluster, pool_id) (ceph_scrub_overdue_pgs{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"})) > 0))\nor (0 * sum(max by (cluster, pool_id) (ceph_scrub_overdue_pgs{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"})))"
}
],
"options": {
@@ -116,31 +116,113 @@
"0": {
"text": "On track",
"color": "green",
"index": 1
}
}
"index": 0
},
{
"type": "range",
"options": {
"from": 0,
"to": 0.01,
"result": {
"text": "Slipping",
"1": {
"text": "Work outstanding",
"color": "blue",
"index": 1
},
"2": {
"text": "Falling behind",
"color": "orange",
"index": 2
},
"3": {
"text": "Deep scrub paused",
"color": "orange",
"index": 3
},
"4": {
"text": "Past warn deadline",
"color": "red",
"index": 4
}
}
}
],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
},
{
"color": "blue",
"value": 1
},
{
"color": "orange",
"value": 2
},
{
"type": "range",
"options": {
"from": 0.01,
"to": null,
"result": {
"text": "Behind",
"color": "red",
"index": 3
"value": 4
}
]
}
},
"overrides": []
},
"description": "Deep-scrub verdict, worst state first. Past warn deadline: a PG is older than the pool's deep_scrub_interval (or the mgr's osd_deep_scrub_interval) x (1 + the mgr's mon_warn_pg_not_deep_scrubbed_ratio). The mgr computes PG_NOT_DEEP_SCRUBBED, so this state mirrors that health check while monk parses every stamp and its config read succeeds. Deep scrub paused: nodeep-scrub is set, which explains every state below it. Falling behind: a PG is more than a day past ceph's latest target, the latest age at which ceph's randomized schedule could have made it eligible (deep interval x (1 + 2 x osd_deep_scrub_interval_cv), 39.2 days on spice). Past that point the PG is late whatever the cause; the day of slack absorbs normal queue wait. Work outstanding: PGs sit past the target interval, but none is more than a day past its latest target. That is normal, not a problem: ceph randomizes deep eligibility around the interval, so a healthy cluster always carries some. On track: nothing is past the target interval. No data: monk's series are absent."
},
{
"id": 34,
"type": "stat",
"title": "Past ceph's latest target (deep)",
"datasource": {
"type": "prometheus",
"uid": "$datasource"
},
"gridPos": {
"h": 3,
"w": 5,
"x": 14,
"y": 4
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "$datasource"
},
"refId": "A",
"instant": false,
"range": true,
"legendFormat": "",
"expr": "max(max by (cluster, pool_id) (ceph_scrub_late_max_seconds{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"}))"
}
],
"options": {
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"orientation": "auto",
"textMode": "value",
"colorMode": "value",
"graphMode": "area",
"justifyMode": "center",
"wideLayout": true
},
"fieldConfig": {
"defaults": {
"unit": "s",
"color": {
"mode": "thresholds"
},
"mappings": [
{
"type": "value",
"options": {
"0": {
"text": "None",
"color": "green",
"index": 0
}
}
}
@@ -154,18 +236,92 @@
},
{
"color": "orange",
"value": 0.0001
},
{
"color": "red",
"value": 0.01
"value": 86400
}
]
}
},
"overrides": []
},
"description": "Plain-language verdict for the deep-scrub cycle, driven by the share of PGs whose last deep scrub is past its pool's target interval. It counts PGs rather than bytes deliberately: pools that store their data in omap, such as the RGW bucket index, report zero bytes and are invisible to the byte-weighted coverage stat. Ceph only makes a PG eligible for deep scrub once it reaches the interval, so the oldest PG's age sits just under the target in normal operation; keying off overdue PGs instead means On track is a state a healthy cluster actually reaches."
"description": "How far the latest PG sits past ceph's latest target: the pool's deep_scrub_interval (or osd_deep_scrub_interval) x (1 + 2 x osd_deep_scrub_interval_cv), the latest date ceph's randomized schedule could have made the PG eligible, 28 days x 1.4 = 39.2 days on spice. A PG past it is late whatever the cause, so this reads None on a healthy cluster apart from brief queue wait; past one day it turns orange and the status block reads Falling behind. PGs whose stamp monk cannot parse are excluded, since lateness needs an established age; PGs overdue counts them instead."
},
{
"id": 35,
"type": "stat",
"title": "Past warn deadline (deep)",
"datasource": {
"type": "prometheus",
"uid": "$datasource"
},
"gridPos": {
"h": 3,
"w": 5,
"x": 19,
"y": 4
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "$datasource"
},
"refId": "A",
"instant": false,
"range": true,
"legendFormat": "",
"expr": "sum(max by (cluster, pool_id) (ceph_scrub_breach_pgs{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"}))"
}
],
"options": {
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"orientation": "auto",
"textMode": "value",
"colorMode": "value",
"graphMode": "area",
"justifyMode": "center",
"wideLayout": true
},
"fieldConfig": {
"defaults": {
"unit": "none",
"color": {
"mode": "thresholds"
},
"mappings": [
{
"type": "value",
"options": {
"0": {
"text": "None",
"color": "green",
"index": 0
}
}
}
],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
},
{
"color": "red",
"value": 1
}
]
}
},
"overrides": []
},
"description": "PGs older than the mon's deep-scrub warn deadline: the pool's deep_scrub_interval option if set, else the mgr's osd_deep_scrub_interval, x (1 + the mgr's mon_warn_pg_not_deep_scrubbed_ratio). The mgr view matters because the mgr is what computes PG_NOT_DEEP_SCRUBBED. This is the number that means data verification has actually slipped and the one to escalate on. It mirrors the health check while monk parses every stamp and its config read succeeds; a PG whose stamp cannot be parsed is not counted here. A ratio of 0 disables the check in ceph, and monk then exports nothing, so this panel shows No data rather than None."
},
{
"id": 2,
@@ -1296,7 +1452,7 @@
"defaults": {
"unit": "percentunit",
"color": {
"mode": "thresholds",
"mode": "fixed",
"fixedColor": "blue"
},
"mappings": [],
@@ -1304,16 +1460,8 @@
"mode": "absolute",
"steps": [
{
"color": "red",
"color": "blue",
"value": null
},
{
"color": "orange",
"value": 0.7
},
{
"color": "green",
"value": 1
}
]
},
@@ -1321,7 +1469,7 @@
},
"overrides": []
},
"description": "Share of stored bytes whose last deep scrub is inside the target interval, measured from every PG's last_deep_scrub_stamp rather than inferred from read counters. Both sides are stored (logical) bytes, deduped per (cluster, pool_id) because every monk instance reports the same cluster-wide data. Byte-weighted, and stat_sum.num_bytes excludes omap: an omap-only pool like the RGW bucket index can go entirely overdue without moving this number, so read it alongside PGs overdue rather than alone."
"description": "Share of stored bytes whose last deep scrub is inside the target interval, measured from every PG's last_deep_scrub_stamp rather than inferred from read counters. Both sides are stored (logical) bytes, deduped per (cluster, pool_id) because every monk instance reports the same cluster-wide data. Byte-weighted, and stat_sum.num_bytes excludes omap: an omap-only pool like the RGW bucket index can go entirely overdue without moving this number, so read it alongside PGs overdue rather than alone. A healthy cluster sits below 100% by the overdue floor, not at it: ceph randomizes eligibility past the interval and pushes a deferred target's not_before, so read the verdict from Past ceph's latest target and Past warn deadline, not from this coverage number. Deliberately uncolored: the healthy value sits below 100% by the overdue floor, so a green-only-at-100% rule would sit orange permanently; color comes from the verdict panels."
},
{
"id": 18,
@@ -1369,7 +1517,7 @@
"defaults": {
"unit": "bytes",
"color": {
"mode": "thresholds",
"mode": "fixed",
"fixedColor": "blue"
},
"mappings": [
@@ -1398,19 +1546,15 @@
"mode": "absolute",
"steps": [
{
"color": "green",
"color": "blue",
"value": null
},
{
"color": "orange",
"value": 1
}
]
}
},
"overrides": []
},
"description": "Stored bytes in PGs whose last deep scrub is older than the target interval. These are logical bytes excluding omap: do not compare them with raw capacity metrics such as ceph_osd_stat_bytes_used without the per-pool expansion factor ceph_pool_stored_raw / ceph_pool_stored, and do not read None here as proof that nothing is overdue; PGs overdue is the byte-agnostic check."
"description": "Bytes in PGs whose last deep scrub is older than the pool's target interval. Same caveat as the PG count: work outstanding, not a breach. Logical bytes, so never mix it with raw capacity metrics like ceph_osd_stat_bytes_used. Deliberately uncolored for the same reason: a healthy cluster always carries some, so color comes from the verdict panels."
},
{
"id": 19,
@@ -1458,7 +1602,7 @@
"defaults": {
"unit": "short",
"color": {
"mode": "thresholds",
"mode": "fixed",
"fixedColor": "blue"
},
"mappings": [
@@ -1477,19 +1621,15 @@
"mode": "absolute",
"steps": [
{
"color": "green",
"color": "blue",
"value": null
},
{
"color": "orange",
"value": 1
}
]
}
},
"overrides": []
},
"description": "PGs whose last deep scrub is older than the target interval. PGs whose stamp could not be parsed are counted here, so a format change reads as overdue rather than as silently clean."
"description": "PGs whose last deep scrub is older than the pool's target interval. This is workload outstanding, not lateness: ceph randomizes eligibility past the interval and pushes a deferred target's not_before, so a healthy cluster always carries some. Use Past ceph's latest target and Past warn deadline for the verdict; use this one for how much work is queued up. Deliberately uncolored for the same reason: a healthy cluster always carries some, so color comes from the verdict panels."
},
{
"id": 20,
@@ -1537,7 +1677,7 @@
"defaults": {
"unit": "percentunit",
"color": {
"mode": "thresholds",
"mode": "fixed",
"fixedColor": "blue"
},
"mappings": [],
@@ -1545,23 +1685,15 @@
"mode": "absolute",
"steps": [
{
"color": "green",
"color": "blue",
"value": null
},
{
"color": "orange",
"value": 1
},
{
"color": "red",
"value": 1.5
}
]
}
},
"overrides": []
},
"description": "Age of the worst pool's oldest deep-scrub stamp, expressed against that pool's own target interval read from the cluster. 100% means a PG has just reached the interval; above that it is overdue. Stated as a ratio rather than days so the panel is correct on any cluster: spice targets 28 days and sietch 7, and a fixed threshold would be wrong on one of them."
"description": "Age of the worst pool's oldest deep-scrub stamp, expressed against that pool's own target interval read from the cluster. 100% means a PG has just reached the interval; above that it is overdue. Stated as a ratio rather than days so the panel is correct on any cluster: spice targets 28 days and sietch 7, and a fixed threshold would be wrong on one of them. It sits above 100% whenever there is work outstanding, which is normal: ceph randomizes deep eligibility past the interval. Compare it with ceph's latest target, 1 + 2 x osd_deep_scrub_interval_cv of the interval (1.4x on spice), and the mon warn deadline, 1 + mon_warn_pg_not_deep_scrubbed_ratio (1.75x on spice). Deliberately uncolored for the same reason as the stats beside it: color comes from the verdict panels."
},
{
"id": 21,
@@ -1710,7 +1842,7 @@
},
"overrides": []
},
"description": "Coverage over time for both scrub depths. Unlike the read-counter estimate this replaces, it needs no warm-up window: it is true from the first scrape."
"description": "Coverage over time for both scrub depths. Unlike the read-counter estimate this replaces, it needs no warm-up window: it is true from the first scrape. The deep line sits below 100% by the overdue floor on a healthy cluster, not at it: ceph randomizes deep eligibility around the interval and pushes deferred targets, so read the verdict from Past ceph's latest target and Past warn deadline, not from coverage. The regular line is judged against osd_scrub_max_interval, well past ceph's own shallow target, so it holds at 100% when healthy."
},
{
"id": 23,
@@ -1762,6 +1894,28 @@
"refId": "D",
"legendFormat": "target (longest)",
"expr": "max(ceph_scrub_target_interval_seconds{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"})"
},
{
"datasource": {
"type": "prometheus",
"uid": "$datasource"
},
"refId": "E",
"instant": false,
"range": true,
"legendFormat": "mon warn deadline (longest pool)",
"expr": "max(ceph_scrub_warn_interval_seconds{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"})"
},
{
"datasource": {
"type": "prometheus",
"uid": "$datasource"
},
"refId": "F",
"instant": false,
"range": true,
"legendFormat": "ceph's latest target",
"expr": "max(ceph_scrub_latest_target_seconds{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"})"
}
],
"options": {
@@ -1834,10 +1988,251 @@
}
}
]
},
{
"matcher": {
"id": "byName",
"options": "mon warn deadline (longest pool)"
},
"properties": [
{
"id": "color",
"value": {
"mode": "fixed",
"fixedColor": "red"
}
},
{
"id": "custom.lineStyle",
"value": {
"fill": "dash",
"dash": [
10,
10
]
}
}
]
},
"description": "Age of the median and 95th-percentile stored byte since its last deep scrub, against the interval read live from the cluster rather than a hardcoded assumption. The quantiles merge every pool, but the target is per pool, so both bounds are drawn: below the shortest target nothing is overdue, above the longest everything is, and between them the merged view cannot say which pool a byte belongs to. The two lines coincide while every pool shares an interval and separate the moment one is overridden; the overdue stats beside this panel stay authoritative either way. The histogram carries gauge semantics under a histogram type: never apply rate() or increase() to these series, and its _count counts bytes, not events. Buckets are deduped per le across instances, so snapshots taken a refresh apart can briefly invert two adjacent buckets; the engine clamps them monotonic, which nudges the quantile rather than failing the query."
{
"matcher": {
"id": "byName",
"options": "ceph's latest target"
},
"properties": [
{
"id": "color",
"value": {
"mode": "fixed",
"fixedColor": "orange"
}
},
{
"id": "custom.lineStyle",
"value": {
"fill": "dash",
"dash": [
10,
10
]
}
}
]
}
]
},
"description": "Age of the median and 95th-percentile stored byte since its last deep scrub, against the interval read live from the cluster rather than a hardcoded assumption. The quantiles merge every pool, but the target is per pool, so both bounds are drawn: below the shortest target nothing is overdue, above the longest everything is, and between them the merged view cannot say which pool a byte belongs to. The two lines coincide while every pool shares an interval and separate the moment one is overridden; the overdue stats beside this panel stay authoritative either way. The histogram carries gauge semantics under a histogram type: never apply rate() or increase() to these series, and its _count counts bytes, not events. Buckets are deduped per le across instances, so snapshots taken a refresh apart can briefly invert two adjacent buckets; the engine clamps them monotonic, which nudges the quantile rather than failing the query. Two dashed lines sit above the targets, both for the longest-interval pool. Orange is ceph's latest target, interval x (1 + 2 x osd_deep_scrub_interval_cv), 39.2 days on spice: the latest date ceph's randomized schedule could make a PG eligible, so only brief queue wait should carry a PG past it. Red is the mon warn deadline, interval x (1 + mon_warn_pg_not_deep_scrubbed_ratio) from the mgr's config, where PG_NOT_DEEP_SCRUBBED fires; it is absent when that ratio is 0 and the check is disabled. These are byte quantiles, so a single late PG can sit past either line without moving p95; the stats in the top row are per PG."
},
{
"id": 36,
"type": "timeseries",
"title": "Backlog composition (deep)",
"datasource": {
"type": "prometheus",
"uid": "$datasource"
},
"gridPos": {
"h": 8,
"w": 8,
"x": 0,
"y": 55
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "$datasource"
},
"refId": "A",
"instant": false,
"range": true,
"legendFormat": "past target interval",
"expr": "sum(max by (cluster, pool_id) (ceph_scrub_overdue_pgs{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"}))"
},
{
"datasource": {
"type": "prometheus",
"uid": "$datasource"
},
"refId": "B",
"instant": false,
"range": true,
"legendFormat": "past ceph's latest target",
"expr": "sum(max by (cluster, pool_id) (ceph_scrub_late_pgs{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"}))"
},
{
"datasource": {
"type": "prometheus",
"uid": "$datasource"
},
"refId": "C",
"instant": false,
"range": true,
"legendFormat": "past the mon warn deadline",
"expr": "sum(max by (cluster, pool_id) (ceph_scrub_breach_pgs{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"}))"
}
],
"options": {
"legend": {
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"fieldConfig": {
"defaults": {
"unit": "none",
"min": 0,
"color": {
"mode": "palette-classic"
},
"custom": {
"drawStyle": "line",
"lineWidth": 2,
"fillOpacity": 10,
"showPoints": "never"
}
},
"overrides": [
{
"matcher": {
"id": "byName",
"options": "past target interval"
},
"properties": [
{
"id": "color",
"value": {
"mode": "fixed",
"fixedColor": "blue"
}
}
]
},
{
"matcher": {
"id": "byName",
"options": "past ceph's latest target"
},
"properties": [
{
"id": "color",
"value": {
"mode": "fixed",
"fixedColor": "orange"
}
}
]
},
{
"matcher": {
"id": "byName",
"options": "past the mon warn deadline"
},
"properties": [
{
"id": "color",
"value": {
"mode": "fixed",
"fixedColor": "red"
}
}
]
}
]
},
"description": "The three thresholds together, which is how to tell a backlog that is draining from one that is growing. Blue is PGs past the target interval: workload outstanding, with a nonzero floor on a healthy cluster. Orange is PGs past ceph's latest target, the latest date its randomized schedule could have made them eligible. Red is PGs past the mon warn deadline, the policy miss PG_NOT_DEEP_SCRUBBED reports. During a scrub pause blue climbs first, orange rises only once PGs pass their latest target, and red once they pass the mon deadline; all of it drains once the flag clears, and the pause annotations mark the window."
},
{
"id": 37,
"type": "timeseries",
"title": "Deep scrub pace",
"datasource": {
"type": "prometheus",
"uid": "$datasource"
},
"gridPos": {
"h": 8,
"w": 8,
"x": 8,
"y": 55
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "$datasource"
},
"refId": "A",
"instant": false,
"range": true,
"legendFormat": "completed per day",
"expr": "sum(max by (cluster, pool_id) (rate(ceph_scrub_completions_total{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"}[6h]))) * 86400"
},
{
"datasource": {
"type": "prometheus",
"uid": "$datasource"
},
"refId": "B",
"instant": false,
"range": true,
"legendFormat": "required per day",
"expr": "sum(max by (cluster, pool_id) (ceph_scrub_pool_pgs{project=\"yucca\", cluster=~\"$cluster\"}) / on (cluster, pool_id) max by (cluster, pool_id) (ceph_scrub_target_interval_seconds{project=\"yucca\", cluster=~\"$cluster\", depth=\"deep\"})) * 86400"
}
],
"options": {
"legend": {
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"fieldConfig": {
"defaults": {
"unit": "short",
"min": 0,
"color": {
"mode": "palette-classic"
},
"custom": {
"drawStyle": "line",
"lineWidth": 2,
"fillOpacity": 0,
"showPoints": "never"
}
},
"overrides": []
},
"description": "Deep scrubs completed per day, from a 6h rate, against the pace needed to keep up: each pool's PG count divided by its deep target interval. Informational, not an input to the verdict. Completed running below required while the backlog beside it grows is the early sign of the scheduler falling behind, well before any PG passes ceph's latest target; short dips are normal because scrubs finish in bursts. The completion counter starts when monk starts, so the line has no history before that and under-reads for the first 6h after a monk restart."
},
{
"id": 24,