Metrics query language reference
This page lists every function, operator, and selector the Uptrace metrics query language accepts. To learn how to write queries, start with Querying metrics; to adapt existing Prometheus queries, see PromQL compatibility.
Not every function works with every metric. Which aggregation a metric accepts depends on its instrument, so read that section before the function tables.
Selectors
A query references a metric through an alias that starts with $:
metrics:
- system_filesystem_usage as $fs_usage
query:
- sum($fs_usage)
Attribute filters
Filter a selector by attributes inside curly braces:
| Operator | Example | Description |
|---|---|---|
= | $fs_usage{state="used"} | Attribute equals the value |
!= | $fs_usage{state!="free"} | Attribute does not equal |
~ | $fs_usage{host_name~"^prod-"} | Attribute matches the regexp |
!~ | $fs_usage{host_name!~"^test-"} | Attribute does not match |
You can also filter every expression in a query at once with where, which supports =, !=, <, <=, >, >=, ~, !~, like, not like, in, and not in:
$hits | $misses | where host_name = 'localhost'
Value filter
The _value pseudo-attribute filters timeseries by the datapoint value, not by an attribute. Use it to count series in a given state:
uniq($status{_value=1}) as num_up | uniq($status{_value=0}) as num_down
Time qualifier
$metric._time reads the datapoint timestamp instead of the value. It accepts only min and max, which return the first and the last time the metric reported:
min($cache._time) | max($cache._time)
Offset
offset shifts the query window. A negative offset looks ahead of the evaluation time:
$http_requests_total offset 5m
$http_requests_total offset -5m
Lookbehind window
Rollup functions read a lookbehind window you can set in square brackets, where i is the current grouping interval:
rate($metric[5i])
max_over_time($metric[1d])
When you omit the window, a rollup reads only the current interval. rate and irate are the exception: they look back max(5 × interval, 5m), so the first buckets of a fine grid still have an earlier sample to rate against.
Instruments
A metric's instrument decides which aggregations it accepts. An aggregation the instrument cannot compute fails the query with <instrument> instrument does not support "<func>".
| Aggregation | Counter | Gauge | PromCounter | Additive | Summary | Histogram |
|---|---|---|---|---|---|---|
sum | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
avg | ✓ | ✓ | ✓ | ✓ | ✓ | |
min, max | ✓ | ✓ | ✓ | ✓ | ✓ | |
median | ✓ | ✓ | ✓ | |||
stddev, stdvar | ✓ | ✓ | ||||
histogram_count | ✓ | ✓ | ||||
p50…p99 | ✓ | |||||
histogram_quantile | ✓ |
- Counter measures a value that only grows, such as the number of processed requests.
min,max, andavgof a growing total carry no meaning, so the instrument rejects them. - Gauge measures a value that goes up and down, such as memory utilization.
sumexists for Prometheus and AWS compatibility, and it double-counts a gauge that reports more than once per interval. - PromCounter is a cumulative counter from Prometheus remote write or the Prometheus receiver. It keeps its running total, so wrap it in
rateorincreaseto read a rate. - Additive measures a value that grows and shrinks and stays meaningful when summed across attributes, such as open connections.
- Summary stores the
min,max,sum, andcountof observed values. It keeps no buckets, so it has no percentiles. - Histogram stores buckets, so it answers percentiles and
histogram_quantile.
uniq and count work with every instrument, because they count timeseries rather than values.
Bare selector
A selector with no aggregation gets a default one, which depends on the instrument:
| Instrument | $metric is the same as |
|---|---|
| Counter, Additive | sum($metric) |
| Gauge, PromCounter | avg($metric) |
Where an aggregation runs
Uptrace pushes an aggregation into ClickHouse when you apply it directly to a selector, which is why sum($metric) scans less data than sum($metric + 0).
count, histogram_count, histogram_quantile, p50…p99, and apdex are pushdown-only: they have no in-process form. Applied anywhere but directly to a metric, they fail with <func>() must be applied to a metric:
# Valid.
p95($srv_duration)
# Invalid — the argument is an expression, not a metric.
p95($a + $b)
Empty buckets
A time bucket with no datapoints reads as absent, and a chart draws a gap. count is the exception: it reads as 0, which makes it suitable for "how many series are reporting" queries.
Aggregate functions
Aggregate functions combine timeseries at each timestamp. Grouping decides which series combine — see Grouping.
| Function | Description |
|---|---|
sum($metric) | Sum of the timeseries. |
avg($metric) | Average of the timeseries. For a histogram, the same as sum($metric) / histogram_count($metric). |
min($metric) | Smallest value. For a histogram or a summary, the smallest observed value. |
max($metric) | Largest value. For a histogram or a summary, the largest observed value. |
median($metric) | Median of the timeseries. |
stddev($metric) | Population standard deviation across the timeseries. Counter and Gauge only. |
stdvar($metric) | Population variance across the timeseries. Counter and Gauge only. |
mad($metric) | Median absolute deviation. Runs in process, so it is not pushed into ClickHouse. |
count($metric) | Number of timeseries that reported. Takes no attributes, and reads 0 for an empty bucket. |
uniq($metric[, attr…]) | Number of distinct values of the listed attributes. With no attributes, the number of timeseries. |
histogram_count($metric) | Number of observations. Summary and Histogram only. |
histogram_quantile(q, $metric) | Quantile q in the range [0..1]. Histogram only. |
p50($metric) | 50th percentile. Histogram only. Also p75, p90, p95, p99. |
apdex($metric, t1, t2) | Apdex score with a satisfied threshold t1 and a tolerating threshold t2. Works with the built-in tracing metrics, which score span duration. |
stddev and stdvar measure how far the timeseries spread apart. On a Counter they measure the spread of per-interval deltas, and on a Gauge the spread of the latest readings. The other instruments reject them: a PromCounter keeps a growing total, an Additive holds complementary parts of one whole, and a Summary or a Histogram no longer holds the individual observations.
Counting timeseries
uniq counts distinct values, and it accepts grouping and a value filter:
# Number of timeseries.
uniq($status) as num_checks
# Distinct combinations of two attributes.
uniq($hits, host_name, service_name) as num_timeseries
# Distinct host_name for each service_name.
uniq($hits by (service_name), host_name) as num_timeseries
Rollup functions
Rollup functions read the datapoints of one timeseries inside the lookbehind window. The number of timeseries stays the same.
| Function | Description |
|---|---|
rate($metric) | Per-second rate of increase. Sizes its window from the grouping interval. |
irate($metric) | Same as rate in the current release. |
increase($metric) | Increase over the window. |
delta($metric) | Same as increase. |
deriv($metric) | Per-second derivative of the values. |
changes($metric) | Number of times the value changed. |
resets($metric) | Number of counter resets. |
min_over_time($metric) | Smallest value in the window. |
max_over_time($metric) | Largest value in the window. |
sum_over_time($metric) | Sum of the values in the window. |
avg_over_time($metric) | Average of the values in the window. |
median_over_time($metric) | Median of the values in the window. |
mad_over_time($metric) | Median absolute deviation in the window. |
stddev_over_time($metric) | Population standard deviation in the window. |
stdvar_over_time($metric) | Population variance in the window. |
first_over_time($metric) | First value in the window. |
last_over_time($metric) | Last value in the window. |
present_over_time($metric) | 1 when the window holds a datapoint. |
absent_over_time($metric) | 1 when the window holds no datapoint. |
default_rollup($metric) | The fill Uptrace applies when a query names no rollup. You rarely write it. |
Use rate or increase to read a PromCounter, which stores a cumulative total:
rate($http_requests_total)
increase($http_requests_total[1h])
Transform functions
Transform functions run on each datapoint of each timeseries. The number of timeseries and the number of datapoints stay the same.
Rate units
| Function | Description |
|---|---|
perMin($metric) | Divides each value by the number of minutes in the grouping interval. |
perSec($metric) | Divides each value by the number of seconds in the grouping interval. |
perMin(sum($cache)) and sum($cache) / _minutes give the same result — see Scalars.
Rounding and sign
| Function | Description |
|---|---|
abs($metric) | Absolute value. |
sgn($metric) | -1, 0, or 1 for the sign of the value. |
ceil($metric) | Rounds up to an integer. |
floor($metric) | Rounds down to an integer. |
trunc($metric) | Drops the fractional part. |
round($metric[, nearest]) | Rounds to the nearest multiple of nearest, which defaults to 1. |
clamp_min($metric, min) | Raises every value below min to min. |
clamp_max($metric, max) | Lowers every value above max to max. |
Exponents and logarithms
| Function | Description |
|---|---|
exp($metric) | e raised to the value. |
exp2($metric) | 2 raised to the value. |
sqrt($metric) | Square root. |
ln($metric) | Natural logarithm. |
log($metric) | Natural logarithm, the same as ln. |
log2($metric) | Base-2 logarithm. |
log10($metric) | Base-10 logarithm. |
Trigonometry
sin, cos, tan, asin, acos, atan, sinh, cosh, tanh, asinh, acosh, and atanh each take one argument and return the named function of every value. deg converts radians to degrees, and rad converts degrees to radians.
Timeseries functions
Timeseries functions rewrite the set of series or fill their gaps. They change how many series a query returns, or which datapoints those series hold.
| Function | Description |
|---|---|
topk($metric, n) | Keeps the n series with the largest values. |
bottomk($metric, n) | Keeps the n series with the smallest values. |
sort($metric) | Orders the series by value, smallest first. |
sort_desc($metric) | Orders the series by value, largest first. |
drop_empty_series($metric) | Removes series that hold no datapoints. |
absent($metric) | Returns 1 for each timestamp the series holds no datapoint. |
timestamp($metric) | Replaces each value with the timestamp of its datapoint. |
clamp($metric, min, max) | Holds every value inside the range. |
interpolate($metric) | Fills a gap by interpolating between the neighbouring datapoints. |
keep_last_value($metric) | Fills a gap with the last known value. |
keep_next_value($metric) | Fills a gap with the next known value. |
# The five busiest hosts.
topk(sum($requests by (host_name)), 5)
# Fill short gaps instead of drawing them.
keep_last_value(avg($temperature))
Operators
Arithmetic and comparison
Uptrace supports +, -, *, /, %, ^, the comparisons ==, !=, <, <=, >, >=, and the set operators and, unless, or, and union.
Precedence runs from highest to lowest:
^*,/,%+,-==,!=,<=,<,>=,>and,unlessor,unionif,ifnot,default
Operators on one level are left-associative, so 2 * 3 % 2 is the same as (2 * 3) % 2. The exception is ^, which is right-associative: 2 ^ 3 ^ 2 is 2 ^ (3 ^ 2).
or keeps a right-side series only where the left side has none with the same attributes. union concatenates both sides without matching attributes at all.
Conditional operators
| Operator | Description |
|---|---|
expr if cond | Keeps a value only where cond has a value. |
expr ifnot cond | Keeps a value only where cond has none. |
expr default val | Replaces an absent value with val. |
Calculate a hit rate only when the sample is large enough, and fill the rest with zero:
(sum($misses) / (sum($hits) + sum($misses))
if (sum($hits) + sum($misses) >= 100)) default 0
The if(cond, then, else) function is deprecated. Use the if, ifnot, and default operators instead.
Grouping and joining
Group at the function level, at the expression level, or for the whole query:
# Function level.
sum($metric) by (host_name, service_name)
avg(sum($metric) by (cpu, mode)) by (cpu)
# Expression level.
sum($metric1) by (type) / sum($metric2) group by host_name
# Query level, affecting every expression.
$metric1 | $metric2 | group by host_name
Math between series joins them by their matching attributes, and one-to-many and many-to-one joins work without extra syntax:
# One-to-one.
$mem_free + $mem_cached group by host_name
# One-to-many.
$cpu_secs by (mode) / $cpu_secs by (service_name, mode)
When the attribute names differ, rename one side:
$metric1 by (hostname as host) + $metric2 by (host_name as host)
Attribute functions
These functions run on attribute values and are valid only in a grouping expression:
| Function | Description |
|---|---|
lower(attr) | Converts the value to lowercase. |
upper(attr) | Converts the value to uppercase. |
trimPrefix(attr, "prefix") | Removes the leading prefix. |
trimSuffix(attr, "suffix") | Removes the trailing suffix. |
extract(attr, pattern) | Extracts the part of the value the regexp captures. |
replace(attr, substring, replacement) | Replaces every occurrence of the substring. |
replaceRegexp(attr, pattern, replacement) | Replaces every part that matches the regexp. |
group by lower(service_name) as service
group by extract(host_name, `^uptrace-prod-(\w+)$`) as host
group by replace(host_name, 'uptrace-prod-', '') as host
Expression aliases
Name an expression with as, then reference the name in a later expression of the same query:
metrics:
- service_cache_redis as $redis
query:
- $redis{type="hits"} as hits
- $redis{type="misses"} as misses
- hits / (hits + misses) as hit_rate
An alias that starts with an underscore takes part in the calculation but does not appear in the result, which keeps helper expressions off the chart:
sum($cache{type="hits"}) as _hits
sum($cache{type="misses"}) as _misses
_misses / (_hits + _misses) as miss_rate
Scalars
| Scalar | Value |
|---|---|
_seconds | Number of seconds in the grouping interval. |
_minutes | Number of minutes in the grouping interval. |
nan | Not-a-number literal. |