Cluster-agent
This page describes metrics gathered by coroot-cluster-agent.
Coroot-cluster-agent is a dedicated tool for collecting cluster-wide telemetry data:
- It gathers database metrics by discovering databases through Coroot's Service Map and Kubernetes control-plane. Using the credentials provided by Coroot or via Kubernetes annotations, the agent connects to the identified databases such as Postgres, MySQL, Redis, Memcached, and MongoDB, collects database-specific metrics, and sends them to Coroot using the Prometheus Remote Write protocol.
- When
--track-database-changesis enabled, the agent tracks schema and configuration changes in databases. Change events are sent to Coroot as OpenTelemetry log records under theDatabaseChangesservice name. - The agent can be integrated with AWS to discover RDS and ElastiCache clusters and collect their telemetry data.
- The agent discovers and scrapes custom metrics from annotated pods.
- The agent monitors GitOps tooling by reading FluxCD and ArgoCD custom resources through its embedded kube-state-metrics and exposing their state as metrics.
- The agent monitors database backups of clusters managed by Kubernetes operators, reading their custom resources through the embedded kube-state-metrics: Postgres (CloudNativePG, Percona Operator for PostgreSQL), MySQL (Percona Operator for MySQL based on Percona XtraDB Cluster), and MongoDB (Percona Operator for MongoDB).
Postgres
pg_up
- Description: Whether the Postgres server is reachable or not
- Type: Gauge
- Source: The agent checks that a connection to the server is still alive on each scrape
pg_probe_seconds
- Description: How long it took to execute an empty SQL query (
;) on the server. This metric shows the round-trip time between the agent and the server - Type: Gauge
- Source: The time spent executing
db.Ping()
pg_scrape_error
- Description: Whether a scrape error occurred
- Type: Gauge
- Labels: error, warning
pg_info
- Description: The server info
- Type: Gauge
- Source:
pg_settings.server_version - Labels: server_version
pg_setting
- Description: Value of the pg_setting variable
- Type: Gauge
- Source:
pg_settings. The agent only collects variables of the following types:integer,realandbool - Labels: name, unit
pg_connections
- Description: The number of the database connections
- Type: Gauge
- Source:
pg_stat_activity - Labels:
- db
- user
- state: current state of the connection, < active | idle | idle in transaction >
- wait_event_type: type of event that the connection is waiting for.
- query - If the state of a connection is
active, this is the currently executing query. Foridle in transactionconnections, this is the last executed query. This label holds a normalized and obfuscated query.
pg_autovacuum_workers
- Description: Number of running autovacuum worker processes
- Type: Gauge
- Source:
pg_stat_activity(backend_type = 'autovacuum worker')
pg_latency_seconds
- Description: Query execution time
- Type: Gauge
- Source:
pg_stat_activity,pg_stat_statements - Labels:
- summary: < avg | max | p50 | p75 | p95 | p99 >
pg_db_queries_per_second
- Description: Number of queries executed in the database
- Type: Gauge
- Source: Aggregation of
pg_stat_activity.state = 'Active'andpg_stat_statements.calls - Labels: db
pg_lock_awaiting_queries
- Description: Number of queries awaiting a lock
- Type: Gauge
- Source: Number of connections with
pg_stat_activity.wait_event_type = 'Lock'. Theblocking_querylabel is calculated using the pg_blocking_pids function - Labels: db, user, blocking_query (the query holding the lock)
Query Metrics
The pg_stat_statements view shows statistics only for queries that have been completed. So, to provide comprehensive statistics, the agent extends this with data about the currently active queries from the pg_stat_activity view.
Collecting stats about each query would produce metrics with very high cardinality. However, the primary purpose of such metrics is to show the most resource-consuming queries. So, the agent collects these metrics only for TOP-20 queries by total execution time.
Each metric described below has query, db and user labels.
Query is a normalized and obfuscated query from pg_stat_statements.query, and pg_stat_activity.query.
For example, the following queries:
SELECT * FROM tbl WHERE id='1';
SELECT * FROM tbl WHERE id='2';
will be grouped to
SELECT * FROM tbl WHERE id=?;
pg_top_query_calls_per_second
- Description: Number of times the query has been executed
- Type: Gauge
- Source:
pg_stat_statements.callsandpg_stat_activity.state = 'Active' - Labels: db, user, query
pg_top_query_time_per_second
- Description: Time spent executing the query
- Type: Gauge
- Source:
clock_timestamp()-pg_stat_activity.query_startandpg_stat_statements.total_time - Labels: db, user, query
pg_top_query_io_time_per_second
- Description: Time the query spent awaiting I/O
- Type: Gauge
- Source:
pg_stat_activity.wait_event_type = 'IO',pg_stat_statements.blk_read_timeandpg_stat_statements.blk_write_time - Labels: db, user, query
Replication metrics
pg_wal_receiver_status
- Description: WAL receiver status: 1 if the receiver is connected, otherwise 0
- Type: Gauge
- Source:
pg_stat_wal_receiverandpg_settings[primary_conninfo] - Labels: sender_host, sender_port
pg_wal_replay_paused
- Description: Whether WAL replay paused or not
- Type: Gauge
- Source:
pg_is_wal_replay_paused()orpg_is_xlog_replay_paused()
pg_wal_current_lsn
- Description: Current WAL sequence number
- Type: Counter
- Source:
pg_current_wal_lsn()orpg_current_xlog_location()
pg_wal_receive_lsn
- Description: WAL sequence number that has been received and synced to disk by streaming replication.
- Type: Counter
- Source:
pg_last_wal_receive_lsn()orpg_last_xlog_receive_location()
pg_wal_reply_lsn
- Description: WAL sequence number that has been replayed during recovery
- Type: Counter
- Source:
pg_last_wal_replay_lsn()orpg_last_xlog_replay_location()
WAL size and archiving
pg_wal_size_bytes
- Description: Size of the WAL directory
- Type: Gauge
- Source:
pg_ls_waldir()
pg_replication_slot_retained_wal_bytes
- Description: Amount of WAL retained for the replication slot
- Type: Gauge
- Source:
pg_replication_slots(restart_lsn);wal_statusis reported on Postgres >= 13 - Labels: slot, active, wal_status
pg_wal_archived_segments_total
- Description: Number of WAL files successfully archived
- Type: Counter
- Source:
pg_stat_archiver
pg_wal_archive_failures_total
- Description: Number of failed attempts to archive WAL files
- Type: Counter
- Source:
pg_stat_archiver
pg_wal_archiving_status
- Description: 1 if the last WAL archive attempt succeeded, 0 if it failed
- Type: Gauge
- Source:
pg_stat_archiver
Checkpoint metrics
The agent tracks checkpoint activity from pg_stat_checkpointer (Postgres >= 17) or pg_stat_bgwriter (older versions). The counters are rebased and accumulated over the agent's lifetime.
pg_checkpoints_scheduled_total
- Description: Number of scheduled checkpoints, including skipped ones
- Type: Counter
- Source:
pg_stat_checkpointer(Postgres >= 17) orpg_stat_bgwriter - Labels: type (
timed,requested)
pg_checkpoints_total
- Description: Number of checkpoints that have been completed
- Type: Counter
- Source:
pg_stat_checkpointer(Postgres >= 17) orpg_stat_bgwriter
pg_restartpoints_total
- Description: Number of restartpoints that have been completed on a standby
- Type: Counter
- Source:
pg_stat_checkpointer.restartpoints_done(Postgres >= 17)
pg_buffers_written_total
- Description: Total number of dirty buffers flushed to disk
- Type: Counter
- Source:
pg_stat_checkpointer(Postgres >= 17) orpg_stat_bgwriter - Labels: source (
checkpointer)
pg_time_since_last_checkpoint_seconds
- Description: Seconds since the last checkpoint observed by the agent
- Type: Gauge
- Source: Measured by the agent from the checkpoint completion counter
pg_wal_since_last_checkpoint_bytes
- Description: Amount of WAL written since the last completed checkpoint (to be replayed in the case of a crash)
- Type: Gauge
- Source:
pg_control_checkpoint()(redo_lsn)
Transaction ID age
pg_xid_age
- Description: Transactions since the oldest unfrozen transaction ID (age of
datfrozenxid) - Type: Gauge
- Source:
pg_database - Labels: db
pg_multixact_age
- Description: Multixacts since the oldest unfrozen multixact ID (age of
datminmxid) - Type: Gauge
- Source:
pg_database - Labels: db
pg_oldest_xmin_age
- Description: Age, in transactions, of the oldest transaction ID held back from freezing, by holder
- Type: Gauge
- Source:
pg_stat_activity,pg_replication_slots,pg_prepared_xacts(Postgres >= 10) - Labels: holder (
running_transaction,standby_feedback,replication_slot,prepared_transaction)
pg_transaction_seconds
-
Description: Age of the longest-running transaction, by query
-
Type: Gauge
-
Source:
pg_stat_activity(now() - xact_start) -
Labels:
- db
- user
- query - the statement the transaction is running. For an
idle in transactionsession this is the last statement it executed, which is the code path that opened the transaction. This label holds a normalized and obfuscated query, and is~emptywhen obfuscation leaves no text (for example a statement that is only a comment).
Where
pg_oldest_xmin_agesays that a running transaction is holding the vacuum horizon back, this says which one, so a report can name the statement instead of leaving you to find it inpg_stat_activityby hand. That matters because the database named in a wraparound or bloat report is whichever one holds the oldestdatfrozenxid, which is often not the database the transaction is running in.Only transactions open for at least 10 seconds are reported, and at most the 20 longest query shapes.
Change tracking
When --track-database-changes is enabled, the agent detects and emits change events for:
- Schema changes — The agent periodically snapshots the DDL of every table (columns, constraints, indexes) across all databases. When a table's DDL changes between consecutive snapshots, a change event is emitted with a unified diff. Only schema modifications are tracked (e.g.,
ALTER TABLE, index creation/removal). The snapshot is collected by connecting to each database and queryingpg_catalogandinformation_schema. - Settings changes — The agent snapshots all
pg_settingsvalues each cycle. When a setting changes (e.g., after a configuration reload or restart), a change event is emitted with the diff. Session-level and client-level overrides are excluded.
Each change event includes db.system, db.target, db.name, db_change.object, and db_change.type attributes.
Size and bloat metrics
The agent collects database and table size metrics. For table sizes, only the top 20 largest tables across all databases are reported.
pg_database_size_bytes
- Description: Total size of the database in bytes
- Type: Gauge
- Source:
pg_database_size() - Labels: db
pg_table_size_bytes
- Description: Total size of the table in bytes including indexes and TOAST
- Type: Gauge
- Source:
pg_total_relation_size() - Labels: db, schema, table
pg_table_size_growth_bytes_per_second
- Description: Table size growth rate in bytes per second. Only the top 20 fastest growing tables across all databases are reported. Requires at least two collection cycles to compute.
- Type: Gauge
- Source: Computed from consecutive
pg_total_relation_size()measurements - Labels: db, schema, table
When --track-database-bloat is enabled, the agent also estimates wasted space (bloat) for tables and indexes from planner statistics (pg_class.reltuples/relpages and pg_stats column widths), without scanning table data. TOAST relations are excluded, and only the top tables and indexes by estimated bloat are reported per database. Estimates are approximate and depend on up-to-date ANALYZE/autovacuum statistics.
pg_db_table_bloat_bytes
- Description: Estimated wasted space across all tables of the database
- Type: Gauge
- Source: Estimated from
pg_classandpg_stats - Labels: db
pg_db_index_bloat_bytes
- Description: Estimated wasted space across all indexes of the database
- Type: Gauge
- Source: Estimated from
pg_classandpg_stats - Labels: db
pg_table_bloat_bytes
- Description: Estimated wasted space in the table heap
- Type: Gauge
- Source: Estimated from
pg_classandpg_stats - Labels: db, schema, table
pg_index_bloat_bytes
- Description: Estimated wasted space in the index
- Type: Gauge
- Source: Estimated from
pg_classandpg_stats - Labels: db, schema, table, index
When --track-database-sizes is enabled, the agent also reports dead-row statistics — a leading indicator of autovacuum falling behind, distinct from bloat. The dead/live counts let Coroot compute autovacuum pressure (how many times past its own autovacuum trigger a table sits), and dead bytes provides materiality. All three are emitted from one query with one top-N ranking (by dead bytes), so every reported table carries the complete set.
pg_table_dead_tuple_bytes
- Description: Estimated size of dead tuples not yet reclaimed by vacuum (heap size × dead fraction)
- Type: Gauge
- Source:
pg_stat_user_tables(n_dead_tup,n_live_tup) andpg_relation_size() - Labels: db, schema, table
pg_table_dead_tuples
- Description: Number of dead tuples not yet reclaimed by vacuum
- Type: Gauge
- Source:
pg_stat_user_tables(n_dead_tup) - Labels: db, schema, table
pg_table_live_tuples
- Description: Estimated number of live tuples
- Type: Gauge
- Source:
pg_stat_user_tables(n_live_tup) - Labels: db, schema, table
pg_table_seconds_since_last_autovacuum
- Description: Seconds since the last autovacuum of the table. Not reported for tables that have never been autovacuumed.
- Type: Gauge
- Source:
pg_stat_user_tables(now() - last_autovacuum) - Labels: db, schema, table
pg_table_setting
- Description: Info metric (value always
1) carrying a table's per-table autovacuum and autoanalyze reloption overrides as labels. Reported only for tables that override at least one setting; an unset override is an empty label. Coroot uses the trigger overrides to compute pressure against the table's own vacuum and analyze triggers, and the cost overrides to pinpoint per-table throttling. - Type: Gauge
- Source:
pg_class.reloptions - Labels: db, schema, table, and the override values:
autovacuum_disabled(1whenautovacuum_enabled=false, which disables autoanalyze too),autovacuum_vacuum_scale_factor,autovacuum_vacuum_threshold,autovacuum_vacuum_cost_delay,autovacuum_vacuum_cost_limit,autovacuum_analyze_scale_factor,autovacuum_analyze_threshold
pg_table_vacuum_in_progress
- Description:
1if a vacuum is currently running on the table. Reported only while a vacuum is in progress. Lets Coroot tell a table that has a worker on it (possibly crawling) from one waiting for a free worker. - Type: Gauge
- Source:
pg_stat_progress_vacuum(Postgres >= 9.6) - Labels: db, schema, table
pg_table_vacuum_throttled
- Description:
1if the running vacuum is sleeping on the cost-based delay (VacuumDelaywait event) at scrape time,0otherwise. Reported only while a vacuum is in progress; a high average means the vacuum is being throttled byautovacuum_vacuum_cost_delay/autovacuum_vacuum_cost_limit. - Type: Gauge
- Source:
pg_stat_progress_vacuumjoined withpg_stat_activity(wait_event = 'VacuumDelay') - Labels: db, schema, table
pg_table_mods_since_analyze
- Description: Rows modified (inserted, updated, or deleted) since the table's planner statistics were last analyzed. Unlike dead tuples this includes inserts, so it is collected as its own top-N (ranked by analyze pressure) rather than reusing the dead-tuple set. Coroot divides it by the autoanalyze trigger to compute how stale the statistics are.
- Type: Gauge
- Source:
pg_stat_user_tables(n_mod_since_analyze) - Labels: db, schema, table
pg_table_reltuples
- Description: The planner's estimate of the number of live rows in the table. Used as the denominator of analyze/vacuum pressure. Unlike
n_live_tupit is persisted inpg_class, so it survives a replica promotion (which resets the cumulativepg_stat_*counters) and keeps pressure from spiking to a false value on a freshly promoted primary. - Type: Gauge
- Source:
pg_class.reltuples - Labels: db, schema, table
pg_table_seconds_since_last_analyze
- Description: Seconds since the table's planner statistics were last refreshed by
ANALYZEor autoanalyze. Not reported for tables that have never been analyzed. - Type: Gauge
- Source:
pg_stat_user_tables(now() - greatest(last_analyze, last_autoanalyze)) - Labels: db, schema, table
Postgres backups
Backup state is collected through the agent's embedded kube-state-metrics from the custom resources of CloudNativePG (cnpg) and the Percona Operator for PostgreSQL (pgBackRest, percona). Every metric carries an operator label identifying the source and, via the common namespace/name labels, correlates to the corresponding DatabaseCluster application in Coroot, so a single operator-agnostic set of metrics powers the backup inspection.
pg_backup_target_info
- Description: A configured backup destination. cnpg reports a single object-storage target, and pgBackRest reports one series per repository. Coroot assembles the destination from the explicit
pathor the object-storage sub-fields. - Type: Info
- Source:
Cluster.spec.backup(cnpg),PerconaPGCluster.spec.backups.pgbackrest.repos(Percona) - Labels: operator, method, path, endpoint, s3_bucket, s3_endpoint, gcs_bucket, azure_container, schedule, retention_policy
pg_cluster_status
- Description: A cluster status condition (e.g.
ReadyForBackup,LastBackupSucceeded,ContinuousArchiving,PGBackRestRepoHostReady). Value is 1 for the currently-active series. The reason feeds Coroot's "why backups are failing" hint. - Type: Info
- Source:
.status.conditions - Labels: operator, type, status, reason
pg_backup_last_successful_timestamp_seconds
- Description: Time of the last successful backup, per method. Reported by cnpg. For Percona it is derived from the individual backup runs.
- Type: Gauge
- Source:
Cluster.status.lastSuccessfulBackupByMethod - Labels: operator, method
pg_backup_first_recoverability_point_timestamp_seconds
- Description: The oldest point in time the cluster can be restored to (start of the recovery window), per method.
- Type: Gauge
- Source:
Cluster.status.firstRecoverabilityPointByMethod - Labels: operator, method
pg_backup_last_failed_timestamp_seconds
- Description: Time of the last failed backup.
- Type: Gauge
- Source:
Cluster.status.lastFailedBackup - Labels: operator
pg_backup_schedule_info
- Description: The backup schedule (cron).
- Type: Info
- Source:
ScheduledBackup.spec.schedule(cnpg) - Labels: operator, cluster, schedule
pg_backup_next_scheduled_timestamp_seconds
- Description: When the next scheduled backup is due. Coroot also derives an expected next run from the schedule and the last backup, so an overdue schedule is detected even when the operator stops advancing this value.
- Type: Gauge
- Source:
ScheduledBackup.status.nextScheduleTime(cnpg) - Labels: operator, cluster
pg_backup_info
- Description: An individual backup run (one series per backup object), used to list recent backups. Carries immutable identity only. The run's phase is in
pg_backup_status. - Type: Info
- Source:
Backup(cnpg),PerconaPGBackup(Percona) - Labels: operator, cluster, method, kind, path
pg_backup_status
- Description: The current phase of a backup run (e.g.
Running,Succeeded,Failed,completed). Value is 1 for the currently-active series (a run's phase changes over its lifecycle, so only the live series is used). - Type: Info
- Source:
Backup.status.phase(cnpg),PerconaPGBackup.status.state(Percona) - Labels: operator, status
pg_backup_completed_timestamp_seconds
- Description: When a backup run completed.
- Type: Gauge
- Source:
Backup.status.stoppedAt(cnpg),PerconaPGBackup.status.completed(Percona) - Labels: operator, cluster
MySQL
mysql_up
- Description: Whether the MySQL server is reachable or not
- Type: Gauge
mysql_scrape_error
- Description: Whether a scrape error occurred
- Type: Gauge
- Labels: error, warning
mysql_info
- Description: The server info
- Type: Gauge
- Labels: server_version, server_id, server_uuid
mysql_top_query_calls_per_second
- Description: Number of times the query has been executed
- Type: Gauge
- Source:
performance_schema.events_statements_summary_by_digest - Labels: schema, query
mysql_top_query_time_per_second
- Description: Time spent executing the query
- Type: Gauge
- Source:
performance_schema.events_statements_summary_by_digest - Labels: schema, query
mysql_top_query_lock_time_per_second
- Description: Time the query spent waiting for locks
- Type: Gauge
- Source:
performance_schema.events_statements_summary_by_digest - Labels: schema, query
mysql_top_query_rows_examined_per_second
- Description: Rows examined per second by the query (vs rows returned: a high ratio means a missing index)
- Type: Gauge
- Source:
performance_schema.events_statements_summary_by_digest - Labels: schema, query
mysql_top_query_rows_sent_per_second
- Description: Rows returned per second by the query
- Type: Gauge
- Source:
performance_schema.events_statements_summary_by_digest - Labels: schema, query
Lock waits
The agent samples performance_schema.data_lock_waits on each scrape and reports the current
number of queries blocked on row locks. Unlike mysql_top_query_lock_time_per_second (which is
derived from cumulative counters and only updates once a blocked query completes), these gauges
reflect the live contention at scrape time, which makes them suitable for correlation with client
latency. Waiters are deduplicated by requesting transaction id, so a query blocked by several
holders is counted once. Not collected on MariaDB.
mysql_locked_queries
- Description: Number of queries currently waiting for a lock
- Type: Gauge
- Source:
performance_schema.data_lock_waitsjoined toperformance_schema.threadsandperformance_schema.events_statements_currenton the requesting thread. Thequerylabel is the obfuscated statement digest of the waiting (victim) query. - Labels: schema, query (the query waiting for the lock)
mysql_lock_awaiting_queries
- Description: Number of queries currently awaiting a lock, grouped by the query holding it
- Type: Gauge
- Source:
performance_schema.data_lock_waitsjoined toperformance_schema.threadsandperformance_schema.events_statements_currenton the blocking thread. Theblocking_querylabel is the obfuscated statement digest of the query holding the lock. - Labels: schema, blocking_query (the query holding the lock)
Replication metrics
mysql_replication_io_status
- Description: Whether the replication IO thread is running
- Type: Gauge
- Labels: source_server_id, source_server_uuid, state, last_error
mysql_replication_sql_status
- Description: Whether the replication SQL thread is running
- Type: Gauge
- Labels: source_server_id, source_server_uuid, state, last_error
mysql_replication_lag_seconds
- Description: Seconds behind master
- Type: Gauge
- Labels: source_server_id, source_server_uuid
Connection metrics
mysql_connections_max
- Description: Maximum number of allowed connections
- Type: Gauge
- Source:
SHOW GLOBAL VARIABLES(max_connections)
mysql_connections_current
- Description: Current number of connected threads
- Type: Gauge
- Source:
SHOW GLOBAL STATUS(Threads_connected)
mysql_connections_total
- Description: Total number of connections since server start
- Type: Counter
- Source:
SHOW GLOBAL STATUS(Connections)
mysql_connections_aborted_total
- Description: Total number of aborted connection attempts
- Type: Counter
- Source:
SHOW GLOBAL STATUS(Aborted_connects)
mysql_connection_errors_max_connections_total
- Description: Number of connections refused because max_connections was reached
- Type: Counter
- Source:
SHOW GLOBAL STATUS
Traffic metrics
mysql_traffic_received_bytes_total
- Description: Total bytes received by the server
- Type: Counter
- Source:
SHOW GLOBAL STATUS(Bytes_received)
mysql_traffic_sent_bytes_total
- Description: Total bytes sent by the server
- Type: Counter
- Source:
SHOW GLOBAL STATUS(Bytes_sent)
mysql_queries_total
- Description: Total number of queries executed
- Type: Counter
- Source:
SHOW GLOBAL STATUS(Questions)
mysql_slow_queries_total
- Description: Total number of slow queries
- Type: Counter
- Source:
SHOW GLOBAL STATUS(Slow_queries)
mysql_top_table_io_wait_time_per_second
- Description: Time spent on table I/O operations
- Type: Gauge
- Source:
performance_schema.table_io_waits_summary_by_table - Labels: schema, table, operation
Change tracking
When --track-database-changes is enabled, the agent detects and emits change events for:
- Schema changes — The agent periodically snapshots the DDL of every table (columns, indexes, foreign keys) across all databases by querying
information_schema. When a table's DDL changes between consecutive snapshots, a change event is emitted with a unified diff. - Settings changes — The agent snapshots all
SHOW GLOBAL VARIABLESvalues each cycle. When a variable changes (e.g., after aSET GLOBALor restart), a change event is emitted with the diff.
Each change event includes db.system, db.target, db.name, db_change.object, and db_change.type attributes.
Size metrics
The agent collects database and table size metrics from information_schema.tables. For table sizes, only the top 20 largest tables across all databases are reported.
mysql_database_size_bytes
- Description: Total size of the database in bytes, computed as the sum of
data_length + index_lengthfor all tables - Type: Gauge
- Source:
information_schema.tables - Labels: db
mysql_table_size_bytes
- Description: Total size of the table in bytes (
data_length + index_length) - Type: Gauge
- Source:
information_schema.tables - Labels: db, table
mysql_table_size_growth_bytes_per_second
- Description: Table size growth rate in bytes per second. Only the top 20 fastest growing tables across all databases are reported. Requires at least two collection cycles to compute.
- Type: Gauge
- Source: Computed from consecutive
information_schema.tablesmeasurements - Labels: db, table
Server load metrics
mysql_threads_running
- Description: Number of threads that are not sleeping
- Type: Gauge
- Source:
SHOW GLOBAL STATUS
mysql_created_tmp_disk_tables_total
- Description: Number of internal on-disk temporary tables created (sorts/joins spilling to disk)
- Type: Counter
- Source:
SHOW GLOBAL STATUS
InnoDB buffer pool
The buffer pool is InnoDB's page cache. Page counts are reported as-is; multiply by
mysql_innodb_page_size_bytes to get bytes. The hit rate is
1 - mysql_innodb_buffer_pool_reads_total / mysql_innodb_buffer_pool_read_requests_total: a
falling hit rate means the working set no longer fits and reads are going to disk.
mysql_innodb_buffer_pool_read_requests_total
- Description: Logical read requests to the InnoDB buffer pool
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_buffer_pool_reads_total
- Description: Read requests that could not be satisfied from the buffer pool and went to disk
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_buffer_pool_write_requests_total
- Description: Writes done to the InnoDB buffer pool
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_buffer_pool_pages_total
- Description: Total pages in the InnoDB buffer pool
- Type: Gauge
- Source:
SHOW GLOBAL STATUS
mysql_innodb_buffer_pool_pages_free
- Description: Free pages in the InnoDB buffer pool
- Type: Gauge
- Source:
SHOW GLOBAL STATUS
mysql_innodb_buffer_pool_pages_dirty
- Description: Dirty (modified, not yet flushed) pages in the buffer pool
- Type: Gauge
- Source:
SHOW GLOBAL STATUS
mysql_innodb_buffer_pool_pages_data
- Description: Pages holding data (clean + dirty) in the buffer pool
- Type: Gauge
- Source:
SHOW GLOBAL STATUS
mysql_innodb_buffer_pool_wait_free_total
- Description: Times a write had to wait for a free buffer-pool page (buffer pool pressure)
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_buffer_pool_pages_flushed_total
- Description: Buffer-pool pages flushed to disk
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_page_size_bytes
- Description: InnoDB page size in bytes (converts buffer-pool page counts to bytes)
- Type: Gauge
- Source:
SHOW GLOBAL VARIABLES
InnoDB row operations and transactions
mysql_innodb_rows_read_total
- Description: Rows read by InnoDB
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_rows_inserted_total
- Description: Rows inserted by InnoDB
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_rows_updated_total
- Description: Rows updated by InnoDB
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_rows_deleted_total
- Description: Rows deleted by InnoDB
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_commit_total
- Description: COMMIT statements executed
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_rollback_total
- Description: ROLLBACK statements executed
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_sort_merge_passes_total
- Description: Merge passes for sorts that spilled to disk (sort_buffer_size too small)
- Type: Counter
- Source:
SHOW GLOBAL STATUS
InnoDB I/O and redo log
mysql_innodb_data_reads_total
- Description: InnoDB data read operations (OS reads)
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_data_writes_total
- Description: InnoDB data write operations (OS writes)
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_data_read_bytes_total
- Description: Bytes read by InnoDB
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_data_written_bytes_total
- Description: Bytes written by InnoDB
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_data_fsyncs_total
- Description: InnoDB fsync() operations
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_log_waits_total
- Description: Times a write had to wait for the redo log buffer to be flushed
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_os_log_written_bytes_total
- Description: Bytes written to the InnoDB redo log
- Type: Counter
- Source:
SHOW GLOBAL STATUS
Lock and contention metrics
Row-lock counters come from SHOW GLOBAL STATUS. Deadlocks, lock-wait timeouts and the undo history
list are not exposed there, so they are read from information_schema.INNODB_METRICS (requires
the PROCESS privilege). A growing history list means purge is lagging behind a long-running
transaction.
mysql_innodb_row_lock_waits_total
- Description: Number of times a row lock had to be waited for
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_row_lock_time_seconds_total
- Description: Total time spent waiting for InnoDB row locks, seconds
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_row_lock_current_waits
- Description: Row locks currently being waited for
- Type: Gauge
- Source:
SHOW GLOBAL STATUS
mysql_table_locks_waited_total
- Description: Table-lock requests that had to wait (LOCK TABLES, MyISAM, DDL metadata locks)
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_table_locks_immediate_total
- Description: Table-lock requests granted immediately
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_innodb_deadlocks_total
- Description: Transactions rolled back by InnoDB deadlocks
- Type: Counter
- Source:
information_schema.INNODB_METRICS
mysql_innodb_lock_wait_timeouts_total
- Description: Transactions rolled back after waiting innodb_lock_wait_timeout for a row lock
- Type: Counter
- Source:
information_schema.INNODB_METRICS
mysql_innodb_history_list_length
- Description: Undo records not yet purged (history list length); grows while a long transaction holds them
- Type: Gauge
- Source:
information_schema.INNODB_METRICS
mysql_innodb_transaction_seconds
-
Description: Age of the longest-running active InnoDB transactions, by query shape
-
Type: Gauge
-
Source:
information_schema.innodb_trx -
Labels: query - a normalized and obfuscated query. InnoDB reports no query for a transaction that is idle between statements, so those are labelled
(idle in transaction).Only transactions open for at least 10 seconds are reported, and at most the 20 longest query shapes. The Postgres equivalent is pg_transaction_seconds.
Galera metrics
Collected when the instance is part of a Galera cluster (Percona XtraDB Cluster, MariaDB Galera),
detected by the presence of wsrep_cluster_size in SHOW GLOBAL STATUS. Flow control is reported as a
counter of seconds, so rate() gives the fraction of time writes were paused.
mysql_wsrep_cluster_size
- Description: Number of nodes in the Galera cluster component
- Type: Gauge
- Source:
SHOW GLOBAL STATUS
mysql_wsrep_cluster_status
- Description: Status of the cluster component this node is in (Primary means it has quorum)
- Type: Gauge
- Source:
SHOW GLOBAL STATUS - Labels: status
mysql_wsrep_local_state
- Description: Current Galera node state (1), with the state name in the label
- Type: Gauge
- Source:
SHOW GLOBAL STATUS - Labels: state
mysql_wsrep_ready
- Description: Whether the node can accept queries (1 = ON)
- Type: Gauge
- Source:
SHOW GLOBAL STATUS
mysql_wsrep_connected
- Description: Whether the node is connected to the cluster (1 = ON)
- Type: Gauge
- Source:
SHOW GLOBAL STATUS
mysql_wsrep_flow_control_paused_seconds_total
- Description: Total time writes were paused by flow control
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_wsrep_local_recv_queue
- Description: Current number of write-sets waiting to be applied on this node
- Type: Gauge
- Source:
SHOW GLOBAL STATUS
mysql_wsrep_local_send_queue
- Description: Current number of write-sets waiting to be sent from this node
- Type: Gauge
- Source:
SHOW GLOBAL STATUS
mysql_wsrep_local_cert_failures_total
- Description: Number of transactions that failed certification (write conflicts)
- Type: Counter
- Source:
SHOW GLOBAL STATUS
mysql_wsrep_local_bf_aborts_total
- Description: Number of local transactions aborted by replication (brute-force aborts)
- Type: Counter
- Source:
SHOW GLOBAL STATUS
Group Replication metrics
Collected when group_replication_group_name is set.
mysql_group_replication_member_state
- Description: Group Replication state of this member (1), with the state name in the label
- Type: Gauge
- Source:
performance_schema.replication_group_members - Labels: state
mysql_group_replication_cluster_size
- Description: Number of members in the Group Replication group
- Type: Gauge
- Source:
performance_schema.replication_group_members
mysql_group_replication_members_online
- Description: Number of members currently ONLINE in the group
- Type: Gauge
- Source:
performance_schema.replication_group_members
mysql_group_replication_transactions_in_queue
- Description: Transactions waiting in the certification queue on this member
- Type: Gauge
- Source:
performance_schema.replication_group_member_stats
mysql_group_replication_transactions_remote_in_applier_queue
- Description: Remote transactions waiting to be applied on this member
- Type: Gauge
- Source:
performance_schema.replication_group_member_stats
mysql_group_replication_conflicts_detected_total
- Description: Number of transactions that failed certification (conflicts) on this member
- Type: Counter
- Source:
performance_schema.replication_group_member_stats
Binary log and undo metrics
Binary logs and undo tablespaces grow independently of table data and are a common cause of a full data volume. If binary logging is disabled the binary log query is skipped.
mysql_binlog_size_bytes
- Description: Total size of the binary logs on disk
- Type: Gauge
- Source:
SHOW BINARY LOGS
mysql_binlog_files
- Description: Number of binary log files on disk
- Type: Gauge
- Source:
SHOW BINARY LOGS
mysql_binlog_expire_seconds
- Description: binlog_expire_logs_seconds; 0 means binary logs are never purged automatically
- Type: Gauge
- Source:
SHOW GLOBAL VARIABLES(binlog_expire_logs_seconds, orexpire_logs_dayson older MySQL and MariaDB)
mysql_undo_size_bytes
- Description: Total size of the InnoDB undo tablespaces on disk
- Type: Gauge
- Source:
information_schema.INNODB_TABLESPACES(INNODB_SYS_TABLESPACESon MariaDB)
MySQL backups
When running on Kubernetes, backup state is collected through the agent's embedded kube-state-metrics from the custom resources of the Percona Operator for MySQL based on Percona XtraDB Cluster (PerconaXtraDBCluster, PerconaXtraDBClusterBackup). Every metric carries an operator label (percona) and, via the common namespace/name labels, correlates to the corresponding DatabaseCluster application in Coroot.
mysql_backup_target_info
- Description: A configured backup storage, one series per entry of
spec.backup.storages(themethodlabel is the storage name). Coroot assembles the destination from the S3 bucket and prefix or the Azure container. - Type: Info
- Source:
PerconaXtraDBCluster.spec.backup.storages - Labels: operator, method, type, s3_bucket, s3_endpoint, s3_prefix, azure_container
mysql_backup_schedule_info
- Description: A scheduled backup (one series per entry of
spec.backup.schedule). Coroot uses the schedule to detect overdue backups;methodis the storage the schedule writes to. - Type: Info
- Source:
PerconaXtraDBCluster.spec.backup.schedule - Labels: operator, task, schedule, method
mysql_backup_pitr_info
- Description: Whether point-in-time recovery (continuous binlog upload) is enabled for the cluster.
- Type: Info
- Source:
PerconaXtraDBCluster.spec.backup.pitr - Labels: operator, enabled
mysql_cluster_status
- Description: The cluster state reported by the operator (e.g.
ready,initializing,error). Value is 1 for the currently-active series. - Type: Info
- Source:
PerconaXtraDBCluster.status.state - Labels: operator, status
mysql_backup_info
- Description: An individual backup run (one series per
PerconaXtraDBClusterBackupobject), used to list recent backups. Carries immutable identity only; the run's phase is inmysql_backup_status.clusteris the name of thePerconaXtraDBClusterthe backup belongs to. - Type: Info
- Source:
PerconaXtraDBClusterBackup.spec(pxcCluster,storageName),.status(storage_type,destination) - Labels: operator, cluster, method, kind, path
mysql_backup_status
- Description: The current phase of a backup run (e.g.
Starting,Running,Succeeded,Failed). Value is 1 for the currently-active series (a run's phase changes over its lifecycle, so only the live series is used). - Type: Info
- Source:
PerconaXtraDBClusterBackup.status.state - Labels: operator, status
mysql_backup_completed_timestamp_seconds
- Description: When a backup run completed.
- Type: Gauge
- Source:
PerconaXtraDBClusterBackup.status.completed - Labels: operator
MongoDB
mongo_up
- Description: Whether the MongoDB server is reachable or not
- Type: Gauge
mongo_scrape_error
- Description: Whether a scrape error occurred
- Type: Gauge
- Labels: error, warning
mongo_info
- Description: The server info;
flavorisperconafor Percona Server for MongoDB,mongodbotherwise - Type: Gauge
- Labels: server_version, flavor
mongo_rs_status
- Description: Replica set status: 1 if the member is part of a replica set. The member reports its own role (
rolelabel), which Coroot uses to identify the primary and to derive each secondary's replication lag frommongo_rs_last_applied_timestamp_ms. - Type: Gauge
- Source:
replSetGetStatus(theselfmember) - Labels: rs, role
mongo_rs_last_applied_timestamp_ms
- Description: Timestamp of the member's last applied operation, in milliseconds. Coroot computes replication lag as the primary's value minus each secondary's.
- Type: Gauge
- Source:
replSetGetStatus(optimes.appliedOpTime)
mongo_rs_member_config_info
- Description: Replica set member configuration
- Type: Gauge
- Source:
replSetGetConfig - Labels: rs, member, arbiter, votes
mongo_rs_config_info
- Description: Replica set configuration; Coroot warns when
write_concern_majority_journal_defaultisfalse(acknowledged majority writes may be lost if a majority of members crash simultaneously) - Type: Gauge
- Source:
replSetGetConfig - Labels: rs, write_concern_majority_journal_default
mongo_profiling_level
- Description: Database profiling level (0 - off, 1 - slow operations, 2 - all operations). Coroot suggests enabling profiling when it is off everywhere, since it is the source of per-query statistics.
- Type: Gauge
- Source: the
profilecommand (get-only,profile: -1) - Labels: db
mongo_oplog_window_seconds
- Description: Time span between the oldest and the newest oplog entries
- Type: Gauge
- Source:
local.oplog.rs
mongo_oplog_size_bytes / mongo_oplog_max_size_bytes
- Description: Current size of the oplog data and the configured maximum
- Type: Gauge
- Source:
collStatsonlocal.oplog.rs
mongo_connections_current / mongo_connections_active / mongo_connections_max
- Description: Number of client connections, connections currently executing operations, and the connection limit
- Type: Gauge
- Source:
serverStatus.connections
mongo_connections_created_total
- Description: Total number of connections created
- Type: Counter
mongo_connections_rejected_total
- Description: Total number of rejected connections
- Type: Counter
mongo_connections_by_app
- Description: Number of connections by client application name (top 20)
- Type: Gauge
- Source:
$currentOp - Labels: app
mongo_opcounters_total
- Description: Total number of operations by type (insert, query, update, delete, getmore, command)
- Type: Counter
- Source:
serverStatus.opcounters - Labels: op
mongo_documents_returned_total
- Description: Total number of documents returned by queries; compared with the scanned counters to detect inefficient queries
- Type: Counter
- Source:
serverStatus.metrics.document
mongo_op_latency_seconds_total / mongo_op_latency_ops_total
- Description: Total operation latency and operation count by type (read, write, command). The average latency is
rate(latency) / rate(ops). - Type: Counter
- Source:
serverStatus.opLatencies - Labels: type
mongo_queued_operations
- Description: Number of operations queued waiting for a lock or a WiredTiger ticket
- Type: Gauge
- Source:
serverStatus.globalLock.currentQueue - Labels: type
mongo_wt_tickets_available
- Description: Available WiredTiger concurrency tickets by type (read, write)
- Type: Gauge
- Source:
serverStatus.queues.execution(8.0+) orserverStatus.wiredTiger.concurrentTransactions(7.0 and earlier) - Labels: type
mongo_wt_cache_used_bytes / mongo_wt_cache_dirty_bytes / mongo_wt_cache_max_bytes
- Description: WiredTiger cache usage, dirty bytes, and the configured maximum
- Type: Gauge
- Source:
serverStatus.wiredTiger.cache
mongo_wt_pages_evicted_by_app_threads_total
- Description: Pages evicted from the WiredTiger cache by application threads - a key cache-pressure signal
- Type: Counter
- Source:
serverStatus.wiredTiger.cache(sum ofpages evicted by application threads,page evict attempts by application threads, andmodified page evict attempts by application threads; the latter two are the 8.0+ counters)
mongo_wt_app_threads_evicting_seconds_total
- Description: Total time application threads spent evicting pages instead of serving queries
- Type: Counter
- Source:
serverStatus.wiredTiger.cache(application thread time evicting (usecs))
mongo_wt_cache_bytes_read_into_total
- Description: Bytes read from disk into the WiredTiger cache (cache misses) - a direct signal that the working set does not fit in cache
- Type: Counter
- Source:
serverStatus.wiredTiger.cache
mongo_wt_checkpoints_total / mongo_wt_checkpoint_seconds_total
- Description: Number of WiredTiger checkpoints completed and the cumulative time spent in them. The rate of the seconds counter is the fraction of wall-time spent checkpointing - sustained high values indicate write-stall risk.
- Type: Counter
- Source:
serverStatus.wiredTiger(transactionon 7.0,checkpointon 8.0+)
mongo_wt_journal_bytes_written_total
- Description: Bytes written to the WiredTiger journal (the crash-recovery write-ahead log)
- Type: Counter
- Source:
serverStatus.wiredTiger.log
mongo_wt_journal_bytes_since_checkpoint
- Description: Bytes written to the journal since the last checkpoint - the amount that would be replayed on crash recovery (MongoDB's analog of Postgres's
pg_wal_since_last_checkpoint_bytes) - Type: Gauge
mongo_time_since_last_checkpoint_seconds
- Description: Time since the last WiredTiger checkpoint observed by the agent; rises above the checkpoint interval if checkpoints stall
- Type: Gauge
mongo_scanned_keys_total / mongo_scanned_documents_total
- Description: Total number of index keys / documents scanned by the query executor. Compared with documents returned to detect inefficient queries.
- Type: Counter
- Source:
serverStatus.metrics.queryExecutor
mongo_collection_scans_total
- Description: Total number of queries that performed a collection scan
- Type: Counter
mongo_scan_and_order_total
- Description: Total number of queries that performed an in-memory sort
- Type: Counter
mongo_write_conflicts_total
- Description: Write conflicts (WiredTiger optimistic-concurrency retries); a rising rate indicates contention on hot documents
- Type: Counter
- Source:
serverStatus.metrics.operation
mongo_ttl_deleted_documents_total
- Description: Total number of documents deleted by TTL indexes
- Type: Counter
mongo_flow_control_time_acquiring_seconds_total
- Description: Total time write operations spent acquiring flow control tickets; a non-zero rate means flow control is throttling writes on the primary
- Type: Counter
- Source:
serverStatus.flowControl
Per-query metrics
Collected from the database profiler (system.profile), which requires operationProfiling to be enabled on
the server, so the numbers cover slow (or, with Percona Server for MongoDB's rateLimit, sampled) operations.
Query shapes are normalized (literal values replaced with ?); only the top 20 shapes by total execution time
are reported. Rates are smoothed over a 60-second window.
Unlike the other MongoDB metrics (read from in-memory counters via serverStatus), these have a server-side cost:
with profiling enabled, mongod writes a system.profile document for every captured operation. The agent only
reads new entries incrementally, but the capture overhead is paid by the server — bound it with slowOpThresholdMs
/ mode: slowOp and, on Percona Server for MongoDB, rateLimit sampling. See
Enabling the profiler for the trade-off.
mongo_top_query_calls_per_second / mongo_top_query_time_per_second / mongo_top_query_docs_returned_per_second / mongo_top_query_docs_examined_per_second / mongo_top_query_keys_examined_per_second
- Description: Per-query-shape execution rate, time, and documents/keys examined vs returned
- Type: Gauge
- Source:
system.profile - Labels: db, collection, query
Current operations
mongo_operations_waiting_for_lock
- Description: Number of operations waiting for a lock, by database
- Type: Gauge
- Source:
$currentOp - Labels: db
mongo_long_running_operations
- Description: Number of operations of this query shape that have been running for at least 10s (their concurrency; top 20 shapes by count). Integrated over time this is the operation-time spent on long-running operations.
- Type: Gauge
- Source:
$currentOp - Labels: db, collection, query, plan
mongo_fsync_locked
- Description: 1 if
db.fsyncLock()is holding a global lock on this member (e.g. a filesystem-snapshot backup). On a secondary this blocks the oplog applier, so it is a direct root cause of replication lag. - Type: Gauge
- Source:
$currentOp(an operation withdesc: "fsyncLockWorker"is present)
mongo_prepared_transactions
- Description: Number of transactions currently in the prepared state. A prepared transaction holds locks on secondaries until its commit/abort replicates, so a stuck or long one can block oplog apply.
- Type: Gauge
- Source:
serverStatus.transactions(currentPrepared)
mongo_open_transactions
- Description: Number of open (including prepared) transactions broken down by the client application that owns them - a long or prepared one holds locks and can block oplog apply. The client name is normalized (UUID/pod suffixes stripped) and empty names are reported as
unknown; only the top 20 applications are kept. - Type: Gauge
- Source:
$currentOp(entries with atransactionsub-document) - Labels: app
Cursors
mongo_cursors_open / mongo_cursors_open_no_timeout / mongo_cursors_timed_out_total
- Description: Open cursors, open cursors created with the
noTimeoutoption (a leak indicator), and the total number of cursors that timed out (a sign a client stopped iterating results before exhausting the cursor) - Type: Gauge (
mongo_cursors_open,mongo_cursors_open_no_timeout), Counter (mongo_cursors_timed_out_total) - Source:
serverStatus.metrics.cursor
Replication apply
mongo_repl_apply_ops_total
- Description: Oplog operations applied by this member - the secondary apply throughput; a low rate on a lagging secondary indicates it is apply-bound
- Type: Counter
- Source:
serverStatus.metrics.repl.apply
mongo_repl_buffer_operations / mongo_repl_buffer_bytes
- Description: Number and size of oplog operations buffered on a secondary waiting to be applied; a backed-up buffer explains replication lag
- Type: Gauge
- Source:
serverStatus.metrics.repl.buffer
MongoDB backups
When running on Kubernetes, backup state is collected through the agent's embedded kube-state-metrics from the custom resources of the Percona Operator for MongoDB (PerconaServerMongoDB, PerconaServerMongoDBBackup), which runs backups with Percona Backup for MongoDB (PBM). Every metric carries an operator label (percona) and, via the common namespace/name labels, correlates to the corresponding DatabaseCluster application in Coroot.
mongo_backup_target_info
- Description: A configured backup storage, one series per entry of
spec.backup.storages(themethodlabel is the storage name). Coroot assembles the destination from the S3 bucket and prefix or the Azure container. - Type: Info
- Source:
PerconaServerMongoDB.spec.backup.storages - Labels: operator, method, type, s3_bucket, s3_endpoint, s3_prefix, azure_container
mongo_backup_schedule_info
- Description: A scheduled backup task (one series per entry of
spec.backup.tasks). Coroot uses the schedule of enabled tasks to detect overdue backups;methodis the storage the task writes to andkindislogicalorphysical. - Type: Info
- Source:
PerconaServerMongoDB.spec.backup.tasks - Labels: operator, task, schedule, method, kind, enabled
mongo_backup_pitr_info
- Description: Whether point-in-time recovery (continuous oplog backup) is enabled for the cluster.
- Type: Info
- Source:
PerconaServerMongoDB.spec.backup.pitr - Labels: operator, enabled
mongo_cluster_status
- Description: The cluster state reported by the operator (e.g.
ready,initializing,error). Value is 1 for the currently-active series. - Type: Info
- Source:
PerconaServerMongoDB.status.state - Labels: operator, status
mongo_backup_info
- Description: An individual backup run (one series per
PerconaServerMongoDBBackupobject), used to list recent backups. Carries immutable identity only; the run's phase is inmongo_backup_status.clusteris the name of thePerconaServerMongoDBthe backup belongs to. - Type: Info
- Source:
PerconaServerMongoDBBackup.spec(clusterName,storageName,type),.status.destination - Labels: operator, cluster, method, kind, path
mongo_backup_status
- Description: The current phase of a backup run (e.g.
waiting,requested,running,ready,error). Value is 1 for the currently-active series (a run's phase changes over its lifecycle, so only the live series is used). - Type: Info
- Source:
PerconaServerMongoDBBackup.status.state - Labels: operator, status
mongo_backup_completed_timestamp_seconds
- Description: When a backup run completed.
- Type: Gauge
- Source:
PerconaServerMongoDBBackup.status.completed - Labels: operator
Change tracking
When --track-database-changes is enabled, the agent detects and emits change events for:
- Index changes — MongoDB is schemaless, so instead of DDL the agent tracks indexes per collection. Each cycle it snapshots all indexes (name and key fields) for every collection. When indexes are added, removed, or modified, a change event is emitted with a unified diff.
- Settings changes — The agent snapshots all server parameters via
getParameter: "*"each cycle. When a parameter changes (e.g., aftersetParameteror a restart), a change event is emitted with the diff.
MongoDB system databases (admin, config, local) are always excluded from index tracking.
Each change event includes db.system, db.target, db.name, db_change.object, and db_change.type attributes.
Size metrics
The agent collects database and collection size metrics. For collection sizes, only the top 20 largest collections across all databases are reported. MongoDB system databases (admin, config, local) are always excluded.
mongo_database_size_bytes
- Description: Total size of the database in bytes
- Type: Gauge
- Source:
listDatabasescommand (sizeOnDisk) - Labels: db
mongo_collection_size_bytes
- Description: Total size of the collection in bytes (data + indexes + storage overhead)
- Type: Gauge
- Source:
collStatscommand (totalSize) - Labels: db, collection
mongo_collection_size_growth_bytes_per_second
- Description: Collection size growth rate in bytes per second. Only the top 20 fastest growing collections across all databases are reported. Requires at least two collection cycles to compute.
- Type: Gauge
- Source: Computed from consecutive
$collStatsmeasurements - Labels: db, collection
mongo_collection_storage_size_bytes
- Description: Bytes allocated on disk for documents of the collection
- Type: Gauge
- Source:
$collStats(storageStats.storageSize) - Labels: db, collection
mongo_collection_free_storage_bytes
- Description: Reusable (fragmented) bytes within the allocated collection storage
- Type: Gauge
- Source:
$collStats(storageStats.freeStorageSize) - Labels: db, collection
mongo_collection_documents
- Description: Number of documents in the collection
- Type: Gauge
- Source:
$collStats(storageStats.count) - Labels: db, collection
Redis
Redis metrics are collected by the embedded redis_exporter (in redis-metrics-only mode with latency histograms disabled), so the full metric set and its semantics are described in the exporter's documentation. The metrics Coroot relies on are:
redis_up
- Description: Whether the Redis server is reachable or not
- Type: Gauge
redis_exporter_last_scrape_error
- Description: Whether a scrape error occurred
- Type: Gauge
- Labels: err
redis_instance_info
- Description: The server info; Coroot uses
role(master/slave) to build the replication topology - Type: Gauge
- Labels: redis_version, role, and other fields of
INFO server/INFO replication
redis_commands_total / redis_commands_duration_seconds_total
- Description: Total number of calls and cumulative execution time per command. The rate of the seconds counter divided by the rate of calls is the average command latency.
- Type: Counter
- Source:
INFO commandstats - Labels: cmd
redis_db_keys / redis_db_keys_expiring
- Description: Number of keys and number of keys with a TTL in each logical database
- Type: Gauge
- Source:
INFO keyspace - Labels: db
Memcached
Memcached metrics are collected by the embedded memcached_exporter, so the full metric set is described in the exporter's documentation. The metrics Coroot relies on are:
memcached_up
- Description: Whether the Memcached server is reachable or not
- Type: Gauge
memcached_version
- Description: The server version
- Type: Gauge
- Labels: version
memcached_limit_bytes
- Description: The configured memory limit for item storage (
-m) - Type: Gauge
memcached_items_evicted_total
- Description: Total number of valid items removed from the cache to free memory for new items; a non-zero rate means the cache is undersized for the working set
- Type: Counter
memcached_commands_total
- Description: Total number of commands by type and outcome (
get/hit,get/miss,set,delete, ...); Coroot derives the hit rate from thegethits and misses - Type: Counter
- Labels: command, status
AWS
When the AWS integration is configured, the agent discovers RDS instances and ElastiCache nodes through the AWS API (optionally filtered by tags) and exposes their state. Every RDS metric carries an rds_instance_id label (<region>/<DBInstanceIdentifier>) and every ElastiCache metric an ec_instance_id label (<region>/<CacheClusterId>/<CacheNodeId>), which Coroot uses to match the instances to the applications that connect to them.
aws_discovery_error
- Description: 1 for each distinct AWS API error encountered during the last discovery cycle, 0 when discovery succeeded
- Type: Gauge
- Labels: error
aws_rds_info
- Description: RDS instance info
- Type: Gauge
- Source:
DescribeDBInstances;ipv4is resolved by the agent from the endpoint address - Labels: region, availability_zone, endpoint, ipv4, port, engine, engine_version, instance_type, storage_type, multi_az, secondary_availability_zone, cluster_id, source_instance_id
aws_rds_status
- Description: The status of the RDS instance (e.g.
available,modifying,backing-up) - Type: Gauge
- Labels: status
aws_rds_allocated_storage_gibibytes / aws_rds_storage_autoscaling_threshold_gibibytes / aws_rds_storage_provisioned_iops
- Description: The allocated storage size, the storage autoscaling upper limit (
MaxAllocatedStorage), and the number of provisioned IOPS - Type: Gauge
aws_rds_backup_retention_period_days
- Description: The number of days automated backups are retained
- Type: Gauge
aws_rds_read_replica_info
- Description: One series per read replica of this instance
- Type: Gauge
- Labels: replica_instance_id
aws_rds_log_messages_total
- Description: Number of messages in the instance's logs (
postgres,aurora-postgresql,mysql,mariadbandaurora-mysqlengines) grouped by the automatically extracted repeated pattern - Type: Counter
- Source: the instance's log files, read through the RDS
DownloadDBLogFilePortionAPI - Labels: level, pattern_hash, sample
The following OS-level metrics are read from RDS Enhanced Monitoring (the RDSOSMetrics CloudWatch Logs group) and are only available when Enhanced Monitoring is enabled for the instance:
aws_rds_cpu_cores
- Description: The number of virtual CPUs
- Type: Gauge
aws_rds_cpu_usage_percent
- Description: The percentage of the CPU spent in each mode
- Type: Gauge
- Labels: mode (
user,system,wait,steal,irq,nice,guest)
aws_rds_memory_total_bytes / aws_rds_memory_cached_bytes / aws_rds_memory_free_bytes
- Description: The total amount of memory, the amount used as page cache, and the amount of unassigned memory
- Type: Gauge
aws_rds_io_ops_per_second / aws_rds_io_bytes_per_second
- Description: The number of I/O operations and bytes read or written per second, per device (
aurora-datafor Aurora's network storage) - Type: Gauge
- Labels: device, operation (
read,write)
aws_rds_io_await_seconds / aws_rds_io_util_percent
- Description: The average time to serve an I/O request including queue time, and the percentage of time during which requests were issued to the device
- Type: Gauge
- Labels: device
aws_rds_io_latency_seconds
- Description: The average elapsed time between the submission of an I/O request and its completion (Amazon Aurora only)
- Type: Gauge
- Labels: device, operation
aws_rds_fs_total_bytes / aws_rds_fs_used_bytes
- Description: The size of each file system and the space used by files on it; Coroot uses the
/rdsdbdatamount point for the data volume - Type: Gauge
- Labels: mount_point
aws_rds_net_rx_bytes_per_second / aws_rds_net_tx_bytes_per_second
- Description: The number of bytes received and transmitted per second, per network interface
- Type: Gauge
- Labels: interface
aws_elasticache_info
- Description: ElastiCache node info;
cluster_idis the replication group id when the node belongs to one, the cache cluster id otherwise - Type: Gauge
- Source:
DescribeCacheClusters;ipv4is resolved by the agent from the endpoint address - Labels: region, availability_zone, endpoint, ipv4, port, engine, engine_version, instance_type, cluster_id
aws_elasticache_status
- Description: The status of the ElastiCache node (e.g.
available,creating,rebooting) - Type: Gauge
- Labels: status
FluxCD
The agent's embedded kube-state-metrics reads FluxCD custom resources (source.toolkit.fluxcd.io, kustomize.toolkit.fluxcd.io, helm.toolkit.fluxcd.io, fluxcd.controlplane.io) and exposes their state. This requires the agent's service account to have get/list/watch access to those API groups, which the Coroot Operator grants automatically.
Every metric below also carries uid, name, and namespace labels identifying the source custom resource. The *_info metrics are info-style: their value is always 1 and the useful data is carried in labels. The *_status metrics expose Kubernetes status conditions: one series per condition, with the value 1 when the condition holds (status: "True") and 0 otherwise ("False"/"Unknown").
fluxcd_git_repository_info / fluxcd_oci_repository_info / fluxcd_helm_repository_info
- Description: Information about a
GitRepository/OCIRepository/HelmRepositorysource - Type: Info
- Source: the source
spec - Labels: url, interval, suspended
fluxcd_git_repository_status / fluxcd_oci_repository_status / fluxcd_helm_repository_status
- Description: Status conditions of a
GitRepository/OCIRepository/HelmRepositorysource - Type: Gauge
- Source:
status.conditions[] - Labels: type (condition type, e.g.
Ready), reason
fluxcd_helm_release_info
- Description: Information about a
HelmRelease - Type: Info
- Source: the HelmRelease
spec - Labels: suspended, interval, target_namespace, source_kind, source_name, source_namespace, chart, version, chart_ref_kind, chart_ref_name, chart_ref_namespace
fluxcd_helm_release_status
- Description: Status conditions of a
HelmRelease - Type: Gauge
- Source:
status.conditions[] - Labels: type, reason
fluxcd_helm_chart_info
- Description: Information about a
HelmChart - Type: Info
- Source: the HelmChart
spec - Labels: chart, version, source_kind, source_name, source_namespace, interval, suspended
fluxcd_helm_chart_status
- Description: Status conditions of a
HelmChart - Type: Gauge
- Source:
status.conditions[] - Labels: type, reason
fluxcd_kustomization_info
- Description: Information about a
Kustomization - Type: Info
- Source: the Kustomization
specandstatus - Labels: suspended, interval, path, source_kind, source_name, source_namespace, target_namespace, last_applied_revision, last_attempted_revision
fluxcd_kustomization_status
- Description: Status conditions of a
Kustomization - Type: Gauge
- Source:
status.conditions[] - Labels: type, reason
fluxcd_kustomization_inventory_entry_info
- Description: A resource managed by a
Kustomization(one series per inventory entry) - Type: Info
- Source:
status.inventory.entries[] - Labels: entry_id
fluxcd_kustomization_dependency_info
- Description: A dependency declared by a
Kustomization(one series perdependsOnentry) - Type: Info
- Source:
spec.dependsOn[] - Labels: depends_on_name, depends_on_namespace
fluxcd_resourceset_info
- Description: Information about a
ResourceSet - Type: Info
- Source: the ResourceSet
status - Labels: last_applied_revision
fluxcd_resourceset_status
- Description: Status conditions of a
ResourceSet - Type: Gauge
- Source:
status.conditions[] - Labels: type, reason
fluxcd_resourceset_inventory_entry_info
- Description: A resource managed by a
ResourceSet(one series per inventory entry) - Type: Info
- Source:
status.inventory.entries[] - Labels: entry_id
fluxcd_resourceset_dependency_info
- Description: A dependency declared by a
ResourceSet(one series perdependsOnentry) - Type: Info
- Source:
spec.dependsOn[] - Labels: depends_on_kind, depends_on_name, depends_on_namespace
ArgoCD
The agent's embedded kube-state-metrics reads ArgoCD Application resources (argoproj.io/v1alpha1) and exposes their sync, health, and operation state. This requires the agent's service account to have get/list/watch access to the argoproj.io API group, which the Coroot Operator grants automatically.
Every metric below also carries uid, name, and namespace labels identifying the Application. All of these are info-style metrics: their value is always 1 and the meaningful state is carried in labels (such as sync_status), so a status that ArgoCD doesn't currently report simply has no series.
argocd_application_info
- Description: Information about an
Application - Type: Info
- Source: the Application
specandstatus - Labels: project, source_type, repo, path, chart, target_revision, dest_server, dest_name, dest_namespace, revision
argocd_application_sync_status
- Description: Sync status of an
Application - Type: Info
- Source:
status.sync.status - Labels: sync_status (e.g.
Synced,OutOfSync,Unknown)
argocd_application_health_status
- Description: Health status of an
Application - Type: Info
- Source:
status.health.status - Labels: health_status (e.g.
Healthy,Progressing,Degraded,Suspended,Missing,Unknown)
argocd_application_operation_status
- Description: Phase of the most recent sync operation
- Type: Info
- Source:
status.operationState.phase - Labels: operation_phase (e.g.
Running,Succeeded,Failed,Error,Terminating)
argocd_application_operation_finished_timestamp_seconds
- Description: When the most recent sync operation finished, as a Unix timestamp
- Type: Gauge
- Source:
status.operationState.finishedAt
argocd_application_resource_info
- Description: A resource managed by an
Application(one series per resource) - Type: Info
- Source:
status.resources[] - Labels: resource_group, resource_kind, resource_namespace, resource_name
argocd_application_resource_sync_status
- Description: Sync status of a managed resource
- Type: Info
- Source:
status.resources[].status - Labels: resource_group, resource_kind, resource_namespace, resource_name, sync_status
argocd_application_resource_health_status
- Description: Health status of a managed resource
- Type: Info
- Source:
status.resources[].health.status - Labels: resource_group, resource_kind, resource_namespace, resource_name, health_status
argocd_application_resource_status
- Description: Result for a resource from the most recent sync operation
- Type: Info
- Source:
status.operationState.syncResult.resources[] - Labels: resource_group, resource_kind, resource_namespace, resource_name, status (e.g.
Synced,Pruned,SyncFailed)
GCP
When the GCP integration is configured, the agent discovers Cloud SQL and Memorystore instances through the GCP APIs (optionally filtered by labels) and exposes their state. Every Cloud SQL metric carries a cloudsql_instance_id label (<project>/<instance>) and every Memorystore metric a memorystore_instance_id label (<project>/<region>/<instance>), which Coroot uses to match the instances to the applications that connect to them.
gcp_discovery_error
- Description: 1 for each distinct GCP API error encountered during the last discovery cycle, 0 when discovery succeeded
- Type: Gauge
- Labels: error
gcp_cloudsql_info
- Description: Cloud SQL instance info
- Type: Gauge
- Source: the Cloud SQL Admin API (
instances.list);ipv4is the private IP if the instance has one, otherwise the public one - Labels: project, region, zone, ipv4, port, engine, engine_version, tier, availability_type, connection_name, instance_type (
CLOUD_SQL_INSTANCEorREAD_REPLICA_INSTANCE), primary_instance (the primary of a read replica)
gcp_cloudsql_status
- Description: The state of the Cloud SQL instance (e.g.
RUNNABLE,SUSPENDED,MAINTENANCE) - Type: Gauge
- Labels: status
gcp_cloudsql_cpu_usage_percent / gcp_cloudsql_cpu_cores / gcp_cloudsql_cpu_usage_cores
- Description: CPU utilization as a percentage of the reserved vCPUs (
database/cpu/utilization), the number of reserved vCPUs (database/cpu/reserved_cores), and the CPU time of the database process in cores (database/cpu/usage_timealigned as a rate) - Type: Gauge
- Source: Cloud Monitoring (
cloudsql.googleapis.com/database/*), the latest 1-minute aligned value
gcp_cloudsql_memory_total_bytes / gcp_cloudsql_memory_used_bytes / gcp_cloudsql_memory_components_percent
- Description: the memory quota (
database/memory/quota), the memory usage of the database process including its buffers and cache (database/memory/total_usage), and the quota split into theusage,cacheandfreecomponents in percent (database/memory/components) - Type: Gauge
- Source: Cloud Monitoring, the latest 1-minute aligned value
- Labels: component (
usage,cache,free) forgcp_cloudsql_memory_components_percent
gcp_cloudsql_disk_total_bytes / gcp_cloudsql_disk_used_bytes
- Description: the data disk quota (
database/disk/quota) and usage (database/disk/bytes_used) - Type: Gauge
- Source: Cloud Monitoring, the latest 1-minute aligned value
gcp_cloudsql_io_ops_per_second
- Description: Disk I/O operations per second of the instance
- Type: Gauge
- Source: Cloud Monitoring,
disk/read_ops_countanddisk/write_ops_countaligned as rates - Labels: operation (
read,write)
gcp_cloudsql_io_bytes_per_second
- Description: Disk I/O throughput of the instance
- Type: Gauge
- Source: Cloud Monitoring,
disk/read_bytes_countanddisk/write_bytes_countaligned as rates - Labels: operation (
read,write)
gcp_cloudsql_network_bytes_per_second
- Description: Network throughput of the instance
- Type: Gauge
- Source: Cloud Monitoring,
network/received_bytes_countandnetwork/sent_bytes_countaligned as rates - Labels: direction (
rx,tx)
gcp_cloudsql_log_messages_total
- Description: Number of messages in the instance's logs (
postgresandmysqlengines) grouped by the automatically extracted repeated pattern - Type: Counter
- Source: Cloud Logging, the
cloudsql_databaseresource of the instance - Labels: level, pattern_hash, sample
gcp_memorystore_info
- Description: Memorystore instance info. Redis and Valkey instances have one series, Memcached instances one per node (
memorystore_instance_idis<project>/<region>/<instance>/<node>) - Type: Gauge
- Source: the Memorystore for Redis, Memcached and Valkey APIs (
instances.list) - Labels: project, region, zone, ipv4, port, engine (
redis,memcached,valkey), engine_version, tier, memory_size_gb, instance
gcp_memorystore_status
- Description: The state of the Memorystore instance (
READYfor Redis and Memcached,ACTIVEfor Valkey, or a transitional state) - Type: Gauge
- Labels: status
gcp_memorystore_cpu_usage_percent / gcp_memorystore_cpu_usage_cores / gcp_memorystore_memory_used_bytes / gcp_memorystore_network_bytes_per_second
- Description: OS-level metrics of the node from Cloud Monitoring, what each product publishes: CPU utilization in percent (Valkey:
instance/cpu/average_utilizationof the primaries), CPU time of the engine process in cores (Redis:stats/cpu_utilizationof the primary summed over user and system time, Memcached:node/cpu/usage_timesummed over the modes), memory used by the engine (Redis:stats/memory/usage, Valkey:instance/memory/total_used_memory, Memcached: the used part ofnode/cache_memory) and network throughput (Redis:stats/network_traffic, Memcached:node/received_bytes_countandnode/sent_bytes_count, aligned as rates) - Type: Gauge
- Source: Cloud Monitoring (
redis.googleapis.com/,memorystore.googleapis.com/instance/,memcache.googleapis.com/node/), the latest 1-minute aligned value - Labels: direction (
rx,tx) forgcp_memorystore_network_bytes_per_second
gcp_memorystore_cpu_cores / gcp_memorystore_memory_total_bytes
- Description: The vCPU count and memory capacity of the node. Memory comes from the instance configuration of all three products; vCPUs are reported only for Memcached nodes (
nodeConfig.cpuCount), the Redis and Valkey APIs don't return them - Type: Gauge
OCI
When the OCI integration is configured, the agent discovers the MySQL HeatWave and Database with
PostgreSQL DB systems and the OCI Cache clusters of the configured compartments. Every DB system metric carries an
oci_db_id label and every cache metric an oci_cache_id label: the OCID of the resource, or the instance id for the
standby instances of a PostgreSQL DB system.
oci_discovery_error
- Description: 1 for each distinct OCI API error encountered since the previous discovery cycle (discovery, OCI Monitoring and the log readers), 0 when there was none
- Type: Gauge
- Labels: error
oci_db_info
- Description: DB system info
- Type: Gauge
- Source: the MySQL HeatWave (
ListDbSystems,ListReplicas) and Database with PostgreSQL (ListDbSystems,GetDbSystem,GetPrimaryDbInstance,GetConnectionDetails) APIs - Labels: name, compartment, region, availability_domain (in the node-agent's form, e.g.
us-ashburn-1-ad-1), ipv4, port, engine (mysql,postgres), engine_version, shape, high_availability, primary (the name of the primary for MySQL read replicas and PostgreSQL standby instances, which are reported as separate instances)
oci_db_status
- Description: The lifecycle state of the DB system (e.g.
ACTIVE,UPDATING,INACTIVE) - Type: Gauge
- Labels: status
oci_db_cpu_cores / oci_db_memory_total_bytes
- Description: the vCPUs (2 per OCPU) and the memory of the DB system, from the shape (PostgreSQL) or from OCI Monitoring (
OCPUsAllocated,MemoryAllocatedfor MySQL) - Type: Gauge
oci_db_cpu_usage_percent / oci_db_cpu_usage_cores / oci_db_memory_used_bytes / oci_db_memory_usage_percent
- Description: CPU utilization (
CPUUtilization), CPU usage in vCPUs (MySQL:OCPUsUsed), memory used by the database (MySQL:MemoryUsed) and memory utilization (PostgreSQL:MemoryUtilization) - Type: Gauge
- Source: OCI Monitoring (
oci_mysql_database,oci_postgresql), the latest 1-minute value
oci_db_disk_total_bytes / oci_db_disk_used_bytes
- Description: storage allocated and used (MySQL:
StorageAllocated,StorageUsed; PostgreSQL:UsedStorage) - Type: Gauge
oci_db_io_ops_per_second / oci_db_io_bytes_per_second / oci_db_io_latency_seconds / oci_db_network_bytes_per_second
- Description: disk I/O operations and throughput (
DbVolumeRead/WriteOperations,DbVolumeRead/WriteBytes,Read/WriteIops,Read/WriteThroughput) and network throughput (NetworkReceive/TransmitBytes, MySQL only) - Type: Gauge
- Labels: operation (
read,write), direction (rx,tx) - Note: the I/O latency (
ReadLatency,WriteLatency) is reported for PostgreSQL only
oci_db_log_messages_total
- Description: the number of messages in the DB system's log grouped by the automatically extracted repeated pattern
- Labels:
level,pattern_hash,sample - Source: OCI Logging, the
postgresql_database_logsservice log of the DB system (PostgreSQL), or the server's error log read throughperformance_schema.error_log(MySQL HeatWave, which publishes no logs to OCI Logging)
oci_cache_info
- Description: OCI Cache cluster info
- Type: Gauge
- Source: the OCI Cache API (
ListRedisClusters);ipv4is the primary endpoint - Labels: name, compartment, region, ipv4, port, engine (
valkey,redis), engine_version, node_count, node_memory_gb
oci_cache_status
- Description: The lifecycle state of the cluster (
ACTIVEor a transitional state) - Type: Gauge
- Labels: status
oci_cache_memory_total_bytes / oci_cache_cpu_usage_percent / oci_cache_memory_used_bytes / oci_cache_network_bytes_per_second
- Description: the memory of a node, and from OCI Monitoring (
oci_redis): CPU utilization (CPUUtilization), memory used by the engine (UsedMemory) and network throughput (NetworkBytesIn/Out) - Type: Gauge
- Labels: direction (
rx,tx) for the network metric
oci_cache_log_messages_total
- Description: the number of messages in the cache cluster's engine log grouped by the automatically extracted repeated pattern
- Labels:
level,pattern_hash,sample - Source: OCI Logging, the
oci-cache-engine-logsservice log of the cluster