Flex Consumption Metrics¶
The Flex Consumption plan bills separately for always-ready instances (kept warm to avoid cold starts) and on-demand instances (allocated per demand and scaled to zero). Each category has its own execution count and execution-unit metric, plus a baseline metric for the always-ready reservation. This page maps those metrics to their billing meters.
flowchart TD
A[Flex Consumption App] --> B[Always Ready Instances]
A --> C[On Demand Instances]
B --> D[AlwaysReadyUnits Baseline]
B --> E[AlwaysReadyFunctionExecutionUnits]
C --> F[OnDemandFunctionExecutionUnits]
D --> G[GB-seconds Meters]
E --> G
F --> G Execution Metrics¶
All Flex Consumption execution metrics use Count unit, Total (Sum) aggregation, and the PT1M time grain.
| Metric | REST name | Meaning |
|---|---|---|
| On Demand Function Execution Count | OnDemandFunctionExecutionCount | Executions that ran on on-demand instances |
| Always Ready Function Execution Count | AlwaysReadyFunctionExecutionCount | Executions that ran on always-ready instances |
| On Demand Function Execution Units | OnDemandFunctionExecutionUnits | MB-milliseconds consumed by on-demand executions |
| Always Ready Function Execution Units | AlwaysReadyFunctionExecutionUnits | MB-milliseconds consumed by always-ready executions |
| Always Ready Units | AlwaysReadyUnits | MB-milliseconds of always-ready capacity held, whether or not executing |
AlwaysReadyUnits is a reservation, not execution
AlwaysReadyUnits measures the capacity you reserve to keep instances warm. It accrues even when no function is running, which is why the always-ready baseline appears on your bill during idle periods.
Mapping Metrics to Billing Meters¶
Execution units are reported in MB-milliseconds; divide by 1,024,000 to convert to GB-seconds. Execution units equal the fixed instance memory size (for example 512, 2048, or 4096 MB) multiplied by total execution time in milliseconds.
| Metric | Billing meter | Conversion |
|---|---|---|
OnDemandFunctionExecutionCount | On Demand Total Executions | Direct count |
AlwaysReadyFunctionExecutionCount | Always Ready Total Executions | Direct count |
OnDemandFunctionExecutionUnits | On Demand Execution Time (GB-s) | / 1,024,000 |
AlwaysReadyFunctionExecutionUnits | Always Ready Execution Time (GB-s) | / 1,024,000 |
AlwaysReadyUnits | Always Ready Baseline (GB-s) | / 1,024,000 |
Performance Metrics¶
| Metric | REST name | Unit | Aggregation | Notes |
|---|---|---|---|---|
| Automatic Scaling Instance Count | InstanceCount | Count | Average | Emitted every 30 seconds; see Scaling and Instances |
| Memory Working Set | MemoryWorkingSet | Bytes | Average | Current memory, per instance |
| Average Memory Working Set | AverageMemoryWorkingSet | Bytes | Average | Average memory, per instance |
| CPU Percentage | CpuPercentage | Percent | Average | Per instance |
CpuPercentage availability
CpuPercentage for Flex Consumption is being rolled out gradually and might not be available for apps in all regions yet. Confirm the metric appears in the metrics explorer before building alerts on it.
Aggregation Guidance¶
- Instance count: because
InstanceCountis emitted every 30 seconds and Flex scales quickly, use a short time grain with Minimum to catch scale-to-zero and Maximum to catch bursts. An average smooths over the exact behavior you usually care about. - Execution units: always use Sum. Charting on-demand versus always-ready execution units side by side shows how much traffic your always-ready reservation is actually absorbing — a key input for right-sizing the reservation.
- Cost attribution: compare
AlwaysReadyUnits(baseline held) againstAlwaysReadyFunctionExecutionUnits(baseline actually used). A large gap means the reservation is oversized for the workload.