Metrics query language reference

This page lists every function, operator, and selector the Uptrace metrics query language accepts. To learn how to write queries, start with Querying metrics; to adapt existing Prometheus queries, see PromQL compatibility.

Selectors

A query references a metric through an alias that starts with $:

yaml
metrics:
  - system_filesystem_usage as $fs_usage
query:
  - sum($fs_usage)

Attribute filters

Filter a selector by attributes inside curly braces:

OperatorExampleDescription
=$fs_usage{state="used"}Attribute equals the value
!=$fs_usage{state!="free"}Attribute does not equal
~$fs_usage{host_name~"^prod-"}Attribute matches the regexp
!~$fs_usage{host_name!~"^test-"}Attribute does not match

You can also filter every expression in a query at once with where, which supports =, !=, <, <=, >, >=, ~, !~, like, not like, in, and not in:

shell
$hits | $misses | where host_name = 'localhost'

Value filter

The _value pseudo-attribute filters timeseries by the datapoint value, not by an attribute. Use it to count series in a given state:

shell
uniq($status{_value=1}) as num_up | uniq($status{_value=0}) as num_down

Time qualifier

$metric._time reads the datapoint timestamp instead of the value. It accepts only min and max, which return the first and the last time the metric reported:

shell
min($cache._time) | max($cache._time)

Offset

offset shifts the query window. A negative offset looks ahead of the evaluation time:

shell
$http_requests_total offset 5m
$http_requests_total offset -5m

Lookbehind window

Rollup functions read a lookbehind window you can set in square brackets, where i is the current grouping interval:

shell
rate($metric[5i])
max_over_time($metric[1d])

When you omit the window, a rollup reads only the current interval. rate and irate are the exception: they look back max(5 × interval, 5m), so the first buckets of a fine grid still have an earlier sample to rate against.

Instruments

A metric's instrument decides which aggregations it accepts. An aggregation the instrument cannot compute fails the query with <instrument> instrument does not support "<func>".

AggregationCounterGaugePromCounterAdditiveSummaryHistogram
sum✓✓✓✓✓✓
avg✓✓✓✓✓
min, max✓✓✓✓✓
median✓✓✓
stddev, stdvar✓✓
histogram_count✓✓
p50…p99✓
histogram_quantile✓
  • Counter measures a value that only grows, such as the number of processed requests. min, max, and avg of a growing total carry no meaning, so the instrument rejects them.
  • Gauge measures a value that goes up and down, such as memory utilization. sum exists for Prometheus and AWS compatibility, and it double-counts a gauge that reports more than once per interval.
  • PromCounter is a cumulative counter from Prometheus remote write or the Prometheus receiver. It keeps its running total, so wrap it in rate or increase to read a rate.
  • Additive measures a value that grows and shrinks and stays meaningful when summed across attributes, such as open connections.
  • Summary stores the min, max, sum, and count of observed values. It keeps no buckets, so it has no percentiles.
  • Histogram stores buckets, so it answers percentiles and histogram_quantile.

uniq and count work with every instrument, because they count timeseries rather than values.

Bare selector

A selector with no aggregation gets a default one, which depends on the instrument:

Instrument$metric is the same as
Counter, Additivesum($metric)
Gauge, PromCounteravg($metric)

Where an aggregation runs

Uptrace pushes an aggregation into ClickHouse when you apply it directly to a selector, which is why sum($metric) scans less data than sum($metric + 0).

count, histogram_count, histogram_quantile, p50…p99, and apdex are pushdown-only: they have no in-process form. Applied anywhere but directly to a metric, they fail with <func>() must be applied to a metric:

shell
# Valid.
p95($srv_duration)

# Invalid — the argument is an expression, not a metric.
p95($a + $b)

Empty buckets

A time bucket with no datapoints reads as absent, and a chart draws a gap. count is the exception: it reads as 0, which makes it suitable for "how many series are reporting" queries.

Aggregate functions

Aggregate functions combine timeseries at each timestamp. Grouping decides which series combine — see Grouping.

FunctionDescription
sum($metric)Sum of the timeseries.
avg($metric)Average of the timeseries. For a histogram, the same as sum($metric) / histogram_count($metric).
min($metric)Smallest value. For a histogram or a summary, the smallest observed value.
max($metric)Largest value. For a histogram or a summary, the largest observed value.
median($metric)Median of the timeseries.
stddev($metric)Population standard deviation across the timeseries. Counter and Gauge only.
stdvar($metric)Population variance across the timeseries. Counter and Gauge only.
mad($metric)Median absolute deviation. Runs in process, so it is not pushed into ClickHouse.
count($metric)Number of timeseries that reported. Takes no attributes, and reads 0 for an empty bucket.
uniq($metric[, attr…])Number of distinct values of the listed attributes. With no attributes, the number of timeseries.
histogram_count($metric)Number of observations. Summary and Histogram only.
histogram_quantile(q, $metric)Quantile q in the range [0..1]. Histogram only.
p50($metric)50th percentile. Histogram only. Also p75, p90, p95, p99.
apdex($metric, t1, t2)Apdex score with a satisfied threshold t1 and a tolerating threshold t2. Works with the built-in tracing metrics, which score span duration.

stddev and stdvar measure how far the timeseries spread apart. On a Counter they measure the spread of per-interval deltas, and on a Gauge the spread of the latest readings. The other instruments reject them: a PromCounter keeps a growing total, an Additive holds complementary parts of one whole, and a Summary or a Histogram no longer holds the individual observations.

Counting timeseries

uniq counts distinct values, and it accepts grouping and a value filter:

shell
# Number of timeseries.
uniq($status) as num_checks

# Distinct combinations of two attributes.
uniq($hits, host_name, service_name) as num_timeseries

# Distinct host_name for each service_name.
uniq($hits by (service_name), host_name) as num_timeseries

Rollup functions

Rollup functions read the datapoints of one timeseries inside the lookbehind window. The number of timeseries stays the same.

FunctionDescription
rate($metric)Per-second rate of increase. Sizes its window from the grouping interval.
irate($metric)Same as rate in the current release.
increase($metric)Increase over the window.
delta($metric)Same as increase.
deriv($metric)Per-second derivative of the values.
changes($metric)Number of times the value changed.
resets($metric)Number of counter resets.
min_over_time($metric)Smallest value in the window.
max_over_time($metric)Largest value in the window.
sum_over_time($metric)Sum of the values in the window.
avg_over_time($metric)Average of the values in the window.
median_over_time($metric)Median of the values in the window.
mad_over_time($metric)Median absolute deviation in the window.
stddev_over_time($metric)Population standard deviation in the window.
stdvar_over_time($metric)Population variance in the window.
first_over_time($metric)First value in the window.
last_over_time($metric)Last value in the window.
present_over_time($metric)1 when the window holds a datapoint.
absent_over_time($metric)1 when the window holds no datapoint.
default_rollup($metric)The fill Uptrace applies when a query names no rollup. You rarely write it.

Use rate or increase to read a PromCounter, which stores a cumulative total:

shell
rate($http_requests_total)
increase($http_requests_total[1h])

Transform functions

Transform functions run on each datapoint of each timeseries. The number of timeseries and the number of datapoints stay the same.

Rate units

FunctionDescription
perMin($metric)Divides each value by the number of minutes in the grouping interval.
perSec($metric)Divides each value by the number of seconds in the grouping interval.

perMin(sum($cache)) and sum($cache) / _minutes give the same result — see Scalars.

Rounding and sign

FunctionDescription
abs($metric)Absolute value.
sgn($metric)-1, 0, or 1 for the sign of the value.
ceil($metric)Rounds up to an integer.
floor($metric)Rounds down to an integer.
trunc($metric)Drops the fractional part.
round($metric[, nearest])Rounds to the nearest multiple of nearest, which defaults to 1.
clamp_min($metric, min)Raises every value below min to min.
clamp_max($metric, max)Lowers every value above max to max.

Exponents and logarithms

FunctionDescription
exp($metric)e raised to the value.
exp2($metric)2 raised to the value.
sqrt($metric)Square root.
ln($metric)Natural logarithm.
log($metric)Natural logarithm, the same as ln.
log2($metric)Base-2 logarithm.
log10($metric)Base-10 logarithm.

Trigonometry

sin, cos, tan, asin, acos, atan, sinh, cosh, tanh, asinh, acosh, and atanh each take one argument and return the named function of every value. deg converts radians to degrees, and rad converts degrees to radians.

Timeseries functions

Timeseries functions rewrite the set of series or fill their gaps. They change how many series a query returns, or which datapoints those series hold.

FunctionDescription
topk($metric, n)Keeps the n series with the largest values.
bottomk($metric, n)Keeps the n series with the smallest values.
sort($metric)Orders the series by value, smallest first.
sort_desc($metric)Orders the series by value, largest first.
drop_empty_series($metric)Removes series that hold no datapoints.
absent($metric)Returns 1 for each timestamp the series holds no datapoint.
timestamp($metric)Replaces each value with the timestamp of its datapoint.
clamp($metric, min, max)Holds every value inside the range.
interpolate($metric)Fills a gap by interpolating between the neighbouring datapoints.
keep_last_value($metric)Fills a gap with the last known value.
keep_next_value($metric)Fills a gap with the next known value.
shell
# The five busiest hosts.
topk(sum($requests by (host_name)), 5)

# Fill short gaps instead of drawing them.
keep_last_value(avg($temperature))

Operators

Arithmetic and comparison

Uptrace supports +, -, *, /, %, ^, the comparisons ==, !=, <, <=, >, >=, and the set operators and, unless, or, and union.

Precedence runs from highest to lowest:

  • ^
  • *, /, %
  • +, -
  • ==, !=, <=, <, >=, >
  • and, unless
  • or, union
  • if, ifnot, default

Operators on one level are left-associative, so 2 * 3 % 2 is the same as (2 * 3) % 2. The exception is ^, which is right-associative: 2 ^ 3 ^ 2 is 2 ^ (3 ^ 2).

or keeps a right-side series only where the left side has none with the same attributes. union concatenates both sides without matching attributes at all.

Conditional operators

OperatorDescription
expr if condKeeps a value only where cond has a value.
expr ifnot condKeeps a value only where cond has none.
expr default valReplaces an absent value with val.

Calculate a hit rate only when the sample is large enough, and fill the rest with zero:

mql
(sum($misses) / (sum($hits) + sum($misses))
  if (sum($hits) + sum($misses) >= 100)) default 0

Grouping and joining

Group at the function level, at the expression level, or for the whole query:

shell
# Function level.
sum($metric) by (host_name, service_name)
avg(sum($metric) by (cpu, mode)) by (cpu)

# Expression level.
sum($metric1) by (type) / sum($metric2) group by host_name

# Query level, affecting every expression.
$metric1 | $metric2 | group by host_name

Math between series joins them by their matching attributes, and one-to-many and many-to-one joins work without extra syntax:

shell
# One-to-one.
$mem_free + $mem_cached group by host_name

# One-to-many.
$cpu_secs by (mode) / $cpu_secs by (service_name, mode)

When the attribute names differ, rename one side:

shell
$metric1 by (hostname as host) + $metric2 by (host_name as host)

Attribute functions

These functions run on attribute values and are valid only in a grouping expression:

FunctionDescription
lower(attr)Converts the value to lowercase.
upper(attr)Converts the value to uppercase.
trimPrefix(attr, "prefix")Removes the leading prefix.
trimSuffix(attr, "suffix")Removes the trailing suffix.
extract(attr, pattern)Extracts the part of the value the regexp captures.
replace(attr, substring, replacement)Replaces every occurrence of the substring.
replaceRegexp(attr, pattern, replacement)Replaces every part that matches the regexp.
shell
group by lower(service_name) as service
group by extract(host_name, `^uptrace-prod-(\w+)$`) as host
group by replace(host_name, 'uptrace-prod-', '') as host

Expression aliases

Name an expression with as, then reference the name in a later expression of the same query:

yaml
metrics:
  - service_cache_redis as $redis
query:
  - $redis{type="hits"} as hits
  - $redis{type="misses"} as misses
  - hits / (hits + misses) as hit_rate

An alias that starts with an underscore takes part in the calculation but does not appear in the result, which keeps helper expressions off the chart:

shell
sum($cache{type="hits"}) as _hits
sum($cache{type="misses"}) as _misses
_misses / (_hits + _misses) as miss_rate

Scalars

ScalarValue
_secondsNumber of seconds in the grouping interval.
_minutesNumber of minutes in the grouping interval.
nanNot-a-number literal.

See also