Cisco Catalyst SD-WAN Manager API
A working reference for every API area: what each endpoint family does, how the Manager serves the call, how to implement it, and where the data actually comes from.
The Manager API is a REST (Representational State Transfer) interface for controlling, configuring, and monitoring devices in an overlay. Cisco frames four use cases: provisioning, network visibility, third-party tool integration, and Network as Code. Every path is prefixed with /dataservice, payloads are JSON except a few file uploads, and the calling role needs API access permission.
Where a call stops decides everything
The Manager sits between your tool and the routers and keeps three stores. A call either reads one of those stores or opens a live NETCONF (Network Configuration Protocol) session to the router. Cost, freshness, and pagination all follow from which one you hit.
offset/limit.count/startId.scrollId/count.- Fig 1 — Data paths: three stores and the device
- Fig 2 / 2b — JWT and API-key login
- Fig 3 — Async task polling
- Fig 4 — Inventory read
- Fig 5 — Template attach (4 steps)
- Fig 6 — Config group lifecycle
- Fig 7 — Multicloud gateway
- Fig 8a — Tunnel vs circuit
- Fig 8 — Bulk device state read
- Fig 9 — Statistics export with scrollId
- Fig 10 — Alarm webhook push
- Fig 11 — Real-time NETCONF round trip
- Fig 12 — Admin-tech generation
- Fig 13 — Client retry policy
- Fig 14 — Empty-response decision tree
Authentication and sessions
Three methods are supported. API keys and JWT (JSON Web Token) arrived in 20.18; session cookies remain for backward compatibility. All three require the XSRF (cross-site request forgery) token on every non-GET request.
Which one to use. API key for long-lived service accounts (collectors, ITSM connectors) — no password in the pipeline, no refresh loop. JWT for short-lived scripts and CI/CD jobs where a bounded token life is a feature. Session cookies only for tooling you cannot change. DevNet's own lab standardises on the API key; verify the exact token-endpoint method (GET vs POST) on your Manager release, since DevNet's pages disagree with each other.
| Method | Login | Token / cookie | Lifetime | Logout |
|---|---|---|---|---|
| JWT | POST /jwt/login — JSON body: username, password, optional duration (seconds) | token claim → Authorization: Bearer; csrf claim → X-XSRF-TOKEN | Default 1,800 s; range 1–604,800 s (7 days). Renew with POST /jwt/refresh | None — keep tokens short |
| API key | Generate once in Manager: click your username → My profile → API token → Generate. No login call. | Authorization: Bearer <apikey>; get XSRF with GET /dataservice/client/token (plain-string body) → X-XSRF-TOKEN | Until revoked in the profile; rotate on a schedule | Revoke in profile |
| Session | POST /j_security_check — form-urlencoded j_username, j_password | JSESSIONID cookie; then GET /dataservice/client/token for XSRF | 24 h, 30-min idle; max 100 concurrent, least-recently-used evicted | POST /logout — always |
Implementation
# Python — minimal JWT client with refresh and XSRF handling
import requests, time
class Manager:
def __init__(self, host, user, pw, duration=1800, verify=True):
self.base = f"https://{host}"; self.s = requests.Session(); self.s.verify = verify
self.user, self.pw, self.duration = user, pw, duration
self.login()
def login(self):
r = self.s.post(f"{self.base}/jwt/login",
json={"username": self.user, "password": self.pw, "duration": self.duration})
r.raise_for_status(); body = r.json()
self.token, self.csrf = body["token"], body.get("csrf")
self.exp = time.time() + self.duration - 60
self.s.headers.update({"Authorization": f"Bearer {self.token}",
"X-XSRF-TOKEN": self.csrf, "Content-Type": "application/json"})
def refresh(self):
r = self.s.post(f"{self.base}/jwt/refresh"); r.raise_for_status()
self.token = r.json()["token"]; self.exp = time.time() + self.duration - 60
self.s.headers["Authorization"] = f"Bearer {self.token}"
def call(self, method, path, **kw):
if time.time() > self.exp: self.refresh()
r = self.s.request(method, f"{self.base}/dataservice{path}", **kw)
if r.status_code == 401: self.login(); r = self.s.request(method, f"{self.base}/dataservice{path}", **kw)
r.raise_for_status(); return r.json()Session-method trap. A failed /j_security_check returns HTTP 200 with an HTML login page in the body. Check for an <html> tag before assuming success. Some administrative calls (login, logout) signal success only through the body.
Pagination, rate limits, and async tasks
Pagination follows the data source
| Store | Parameters | Example | Notes |
|---|---|---|---|
| Configuration DB | offset, limit | GET /template/feature?offset=1&limit=10 | Snapshot at call time; no consistency guarantee between calls |
| Statistics DB | scrollId, count | GET /statistics/approute/page?scrollId=…&count=10 | scrollId expires after 10 minutes; loop until hasMoreData is false |
| Device state cache | count, startId | GET /data/device/state/BFDSessions?count=1000 | pageInfo.moreEntries; pass returned endId as next startId |
| Device (real-time) | — | GET /device/bfd/sessions?deviceId=… | Pagination decided by the device |
Sorting and filtering are only available for device state and statistics APIs; the allowed field names for a table come from its /fields endpoint.
Rate limits
| Limit | Value | Where it bites |
|---|---|---|
Bulk API (/data/device/statistics/) | 48 requests/minute per node (from 20.6); from 20.10.1 requests are distributed across a cluster, so effective limit = per-node × node count | Statistics exports |
| Bulk statistics concurrency | 2 concurrent requests | Parallel exporters |
| All other APIs | 100 requests/second | Tight polling loops |
| Concurrent sessions | 250 on the Manager | Many scripts without logout |
| Cisco SD-WAN Cloud (SaaS fabric via API gateway, 20.15+) | Real-time and bulk APIs are not available; other APIs limited to 10/s at the gateway and 5/s at the Manager; only certificate, configuration, inventory, monitoring and troubleshooting categories are exposed | Cisco-operated SaaS only — not Cisco-hosted control components |
The per-node bulk limit is tunable with request nms server-proxy set ratelimit followed by request nms server-proxy restart on every node. Exhausting a limit returns 429 with no Retry-After header, so back off on your own schedule.
Every write is a task
Configuration pushes and device actions return a task ID, not a result. Poll GET /device/action/status/{taskId} and read the activity array until statusId settles.
def wait_task(m, task_id, timeout=900):
t0 = time.time()
while time.time() - t0 < timeout:
st = m.call("GET", f"/device/action/status/{task_id}")
if st["summary"]["status"] == "done":
return [(d["deviceIP"], d["statusId"]) for d in st["data"]]
time.sleep(8)
raise TimeoutError(task_id)RBAC (role-based access control)
Each operation in the OpenAPI spec carries an x-roles-required value such as Template Deploy-write or Config Group > Device > Deploy-write. Build service accounts from those values: a read-only monitoring group for collectors, and a separate Settings-write account only for webhook rule management.
Administration and settings Configuration DB
Global Manager parameters, user and group management, tenants, software maintenance, and backup. Everything here reads and writes the configuration database; nothing touches a device except software install and reboot actions.
| Endpoint | Function | Use case |
|---|---|---|
| GET/POST /admin/user | List and create users | Service accounts for collectors and CI/CD (continuous integration / continuous delivery) |
| GET/POST /admin/usergroup | Groups as lists of features with read/write flags (Alarms, Audit Log, Device Monitoring, Template Deploy…) | Least-privilege roles |
| GET/POST /admin/resourcegroup | Scope users to sites or regions | Regional NOC access |
| GET/POST /settings/configuration/{type} | Org name, Validator address, certificate authorization, statistics collection settings | Day-0 automation |
| POST /device/action/software, /install, /changepartition | Image upload, install, activate, set default | Upgrade pipelines |
| GET/POST /tenant, /tenantbackup | Multitenant operations (Provider view) | MSP (managed service provider) onboarding |
| GET/PUT /statistics/settings/disable/devicelist/{indexName} | Turn statistics collection on or off per index and device list | Reducing Manager load |
How a user is created
- GET /admin/usergroup to confirm the group exists, or POST /admin/usergroup with the feature list.
- POST /admin/user with
userName,password,group[], optionalresGroupName. - Verify with GET /admin/user; the password is never returned.
m.call("POST", "/admin/usergroup", json={
"groupName": "monitoring-ro",
"tasks": [{"feature": "Device Monitoring", "enabled": True, "read": True, "write": False},
{"feature": "Alarms", "enabled": True, "read": True, "write": False},
{"feature": "Device Inventory", "enabled": True, "read": True, "write": False}]})
m.call("POST", "/admin/user", json={"userName": "svc-collector", "password": "…", "group": ["monitoring-ro"]})Device inventory Configuration DB
The foundation for everything else: which devices exist, whether they are reachable, and what state their certificates and control connections are in. All served from the Manager's own database.
| Endpoint | Returns | Use case |
|---|---|---|
| GET /device | All connected devices: system-ip, host-name, reachability, site-id, version, BFD session counts, OMP peer counts, control connections, uptime, coordinates. 26.1 adds a site-id filter. | CMDB (configuration management database) sync; health rollups |
| GET /system/device/controllers | Validators, Controllers, Managers with certificate expiry, config sync state | Certificate watchdog |
| GET /system/device/vedges | All WAN edges (vEdge and Catalyst IOS XE): cert state, configOperationMode (cli / vmanage), validity | Onboarding audits |
| POST /system/device/fileupload | Upload the authorized serial file | ZTP (zero-touch provisioning) pipelines |
| PUT /system/device/{uuid} / /decommission/{uuid} | Set valid/invalid/staging, decommission | Lifecycle |
| GET /system/device/bootstrap/device/{uuid} | Generate bootstrap configuration | Day-0 shipping |
| GET /device/models | Device models supported, interface naming per model | Template validation |
| GET /health/devices/overview | Good / fair / poor counts (new shape in 26.1) | One-call fabric health |
devices = m.call("GET", "/device")["data"]
down = [d for d in devices if d.get("reachability") != "reachable"]
print(f"{len(devices)} devices, {len(down)} unreachable")UX 1.0 — templates and centralized policy Writes config
UX (user experience) 1.0 is the template era: feature templates compose into device templates, which are attached to devices with per-device variables. Centralized policy is assembled and activated on the Controller. Still the majority of brownfield estates.
| Endpoint | Function |
|---|---|
| GET/POST /template/feature, /template/feature/object/{id}, /template/feature/types | Feature templates (system, VPN, interface, BFD…) |
| GET /template/device, POST /template/device/feature, POST /template/device/cli, GET /template/device/object/{id} | Device templates; the object call returns full JSON (clone pattern) |
| POST /template/device/config/input | Generate the variable sheet for templateId + deviceIds |
| POST /template/device/config/config | Preview the rendered configuration |
| POST /template/device/config/attachfeature / /attachcli | Push; returns a task ID |
| POST /template/config/device/mode/cli | Detach (return to CLI mode) |
| /template/policy/list/*, /template/policy/definition/*, /template/policy/vsmart, /template/policy/vedge, /template/policy/security | Lists (site, VPN, prefix, SLA, apps), definitions (control, data, app-route, cflowd, hub-and-spoke, mesh, zone-based firewall), centralized, localized, security policy |
tpl = "a1b2…"; dev = "C8K-…uuid"
sheet = m.call("POST", "/template/device/config/input",
json={"templateId": tpl, "deviceIds": [dev], "isEdited": False, "isMasterEdited": False})
row = sheet["data"][0]
row.update({"//system/host-name": "BR-100-R1", "//system/system-ip": "10.255.1.1", "//system/site-id": "100"})
task = m.call("POST", "/template/device/config/attachfeature",
json={"deviceTemplateList": [{"templateId": tpl, "device": [row], "isEdited": False, "isMasterEdited": False}]})
print(wait_task(m, task["id"]))UX 2.0 — configuration groups and feature profiles Writes config
The strategic direction for new deployments. A configuration group is a bundle of feature profiles (system, transport, service, policy-object, CLI, other); devices are associated to the group, receive per-device variables, and are deployed in one task. Policy groups and topology groups follow the same CRUD (create, read, update, delete) + associate + deploy pattern.
| Endpoint | Function |
|---|---|
| POST/GET /v1/config-group, PUT/DELETE /v1/config-group/{id} | Group lifecycle |
| PUT /v1/config-group/{id}/device/associate | Bind devices |
| PUT /v1/config-group/{id}/device/variables | Per-device variable values |
| POST /v1/config-group/{id}/device/deploy | Deploy; returns task ID; role Config Group > Device > Deploy-write |
| /v1/policy-group, /v1/topology-group | Same pattern for policy and topology |
| /v1/network-hierarchy | Regions and sites that groups reference |
| /v1/feature-profile/sdwan/{system|transport|service|policy-object|cli|other}/{profileId}/{feature} | Parcel-level features, e.g. …/system/{id}/aaa, …/system/{id}/omp, …/transport/{id}/wan/vpn/{vpnId}/interface/ethernet, …/service/{id}/lan/vpn |
| /v1/feature-profile/sd-routing/… | Same model for autonomous (non-SD-WAN) routers, e.g. …/cli/{cliId}/full-config/{fullConfigId} |
gid = m.call("POST", "/v1/config-group", json={"name": "branch-std", "solution": "sdwan", "description": "Standard branch",
"profiles": [{"id": SYSTEM_PROFILE}, {"id": TRANSPORT_PROFILE}, {"id": SERVICE_PROFILE}]})["id"]
m.call("PUT", f"/v1/config-group/{gid}/device/associate", json={"devices": [{"id": dev}]})
m.call("PUT", f"/v1/config-group/{gid}/device/variables",
json={"solution": "sdwan", "devices": [{"device-id": dev,
"variables": [{"name": "host_name", "value": "BR-100-R1"}, {"name": "system_ip", "value": "10.255.1.1"}, {"name": "site_id", "value": 100}]}]})
task = m.call("POST", f"/v1/config-group/{gid}/device/deploy", json={"devices": [{"id": dev}]})
wait_task(m, task["parentTaskId"])SD-WAN services and partner integrations Writes config
SD-WAN services
- /multicloud/* — cloud account onboarding (AWS, Azure, GCP), cloud gateway creation, VPC/VNet discovery and intent mapping.
- /template/cloudx/* — Cloud OnRamp for SaaS application enablement and DIA (direct Internet access) probes.
- /cloudservices/* — cloud service tokens and Microsoft 365 preferred-path settings (26.1 changed the M365 request schema).
- 26.1 deprecates the /dca/* data-collection endpoints.
Partner integrations
- Webex and Cisco Secure Access / SIG (secure Internet gateway) credentials and tunnel objects.
- ThousandEyes Enterprise Agent enablement — a Docker container on IOS XE edges from 17.6.1, configured as a feature-profile parcel (
…/other/{id}/thousandeyes). - SSE (security service edge) tunnel automation.
Customer questions → endpoints Statistics DB
NOCs (network operations centres) don't ask for endpoint families; they ask "is site 90 healthy?" This table maps the common questions to the scalable endpoint and to the troubleshooting-only equivalent. Every scalable row is served from the Manager; none touches a device.
| Question | Scalable endpoint (poll this) | Troubleshooting-only equivalent |
|---|---|---|
| Which devices are at a site, and are they reachable? | GET /device?site-id={id} | — |
| Is each WAN edge healthy (good / fair / poor)? | GET /statistics/devicehealth/overview/cpu?last_n_hours=1&limit=100 (optionally &site=). Manager-calculated score from CPU, memory, QoE, reachability | GET /device/system/status?deviceId= |
| What did CPU, memory and disk look like over time? | POST /statistics/system with entry_time and vdevice_name rules | same |
| How good is each tunnel path (latency / jitter / loss)? | POST /statistics/approute/aggregation grouped by local_system_ip, local_color, remote_system_ip, remote_color | GET /device/app-route/statistics?deviceId= |
| Is a circuit (transport) up right now? | GET /data/device/state/BFDSessions?count= grouped client-side by device + local-color | GET /device/bfd/sessions?deviceId= |
| How good was each circuit historically? | POST /statistics/approute/aggregation grouped by vdevice_name, local_color | — |
| How available was each site (downtime)? | POST /statistics/nwa/details with type=site, sum down_time by site_id (NWA = network availability) | — |
| Composite site health, ready-made? | GET /statistics/sitehealth/common?last_n_hours=24&includeDetails=true | — |
| Circuit availability across the fabric? | POST /statistics/nwa/aggregation with type=link, grouped by system_ip, color | — |
| Which applications are seen at a site? | GET /statistics/dpi/applications?query={url-encoded JSON} with every site edge in one vdevice_name rule | GET /device/dpi/applications?deviceId= |
| Which app families use the most bandwidth? | POST /statistics/dpi/aggregation grouped by family, sum octets | — |
| Application health across sites? | GET /statistics/perfmon/applications/sites/health?last_n_hours=24&includeUsage=true | — |
Tunnel is not circuit
Customers use the words interchangeably; the API doesn't. A tunnel (path) is one local device + colour to one remote device + colour. A circuit (transport) is one local device + colour, carrying many tunnels. Group your aggregation accordingly, and never group by local_color alone across devices — it merges every "mpls" in the fabric into one row.
The same device has four names
Inventory uses hyphenated fields, statistics use underscores, and real-time calls use deviceId. Mixing them up is the number one cause of "the API returns nothing" tickets.
Inventory field (GET /device) | Used as | Where |
|---|---|---|
site-id | site-id query parameter · site_id response field · site parameter on devicehealth | Inventory filter, statistics responses, health overview |
system-ip | deviceId | Every real-time call: /device/…?deviceId= |
system-ip | vdevice_name | Every statistics query rule and response |
system-ip | system_ip / local_system_ip | Health overview, NWA, app-route aggregation |
reachability | Selection condition | Skip real-time calls to unreachable devices |
personality | vedge vs controllers | Exclude controllers from edge-only queries |
Never use a hostname, management IP, or chassis UUID where an operation asks for the system IP. And never build on /analytics/api/v4/dataservice/aggregate/* — those paths are used by the Manager UI but are not in the OpenAPI spec and can change without notice; every result they give has a documented /statistics equivalent above.
Device state — bulk State cache
The "what is up or down right now" view. The Manager's collection manager polls devices on its own schedule and stores current state; you read the cache, and the router is never touched. One request returns the whole fabric in batches.
| Endpoint | Detail |
|---|---|
| GET /data/device/state/{DataType}?count=N | count mandatory (1–10,000); data types are case-sensitive |
| Data types | BFDSessions · BGPNeighbor · Bridge · ControlConnection · ControlLocalProperty · ControlWanInterface · HardwareAlarms · HardwareEnvironment · HardwareInventory · Interface · OMPPeer · SystemStatus · System |
| Paging | Response pageInfo carries moreEntries and endId; call again with startId=endId |
| Per-device cached view | GET /device/{feature}/synced?deviceId= — same shape as real-time, served from the NMS (network management system) cache |
| Sync control | POST /device/blockSync?blockSync=true|false stops or resumes the collection manager — if state looks stale, check this first |
def state(m, table, count=1000):
rows, start = [], None
while True:
q = f"count={count}" + (f"&startId={start}" if start else "")
r = m.call("GET", f"/data/device/state/{table}?{q}")
rows += r["data"]; pi = r.get("pageInfo", {})
if not pi.get("moreEntries"): return rows
start = pi["endId"]
bfd = state(m, "BFDSessions")
down = [s for s in bfd if s.get("state") != "up"]Statistics — export and query Statistics DB
Devices export counters to the Manager, which stores them in a time-series database. Two flavours share the same store: a bulk export for moving raw rows out, and a query/aggregation API that does the maths server-side and powers the Manager GUI charts.
Bulk export
| Item | Detail |
|---|---|
| Endpoint | GET /data/device/statistics/{datatype}?startDate=&endDate=&count=&timeZone= — start and end mandatory, yyyy-MM-ddThh:mm:ss |
| Data types | alarm · approutestatsstatistics (tunnel loss/latency/jitter) · auditlog · cflowdstatistics · cloudxstatistics · deviceconfiguration · deviceevent · devicesystemstatusstatistics (CPU/memory) · dpistatistics · flowlogstatistics · interfacestatistics · wlanclientinfostatistics |
| Helpers | …/{datatype}/doccount, …/{datatype}/fields; v2 interface stats at GET /v2/data/device/statistics/interfacestatistics |
| Paging | scrollId in the response; pass it back until hasMoreData is false; expires after 10 minutes |
| Limits | 2 concurrent, 48/min per node |
Discover fields before you filter
Every statistics table exposes GET /statistics/{table}/fields (fields you can return) and GET /statistics/{table}/query/fields (fields you can filter on). Field names differ between releases; hard-coding them without checking is how integrations break silently on upgrade.
fields = {f["property"] for f in m.call("GET", "/statistics/approute/fields")}
assert {"latency","jitter","loss_percentage","local_color"} <= fieldsQuery and aggregation
| Item | Detail |
|---|---|
| Endpoint family | GET|POST /statistics/{table} plus /aggregation, /csv, /doccount, /fields, /page; 26.1 adds page, pageSize, sortBy parameters |
| Tables | approute · interface · dpi · qos · flowlog · system (/cpu, /memory) · sitehealth · tunnelhealth · devicehealth · perfmon · fwall · urlf · ipsalert · umbrella · speedtest · qfp · qfpdrop · temperature · powerconsumption · endpointTracker · eiolte · wlanclientinfo · art · apphosting · sul · ppl/ctg and ppl/reordering (new in 26.1) |
| Query body | query (condition AND/OR + rules: field, type, value[], operator) · sort (field/order pairs) · fields · aggregation (histogram buckets + metrics) |
| Deprecated in 26.1 | /statistics/bfd /statistics/cflowd /statistics/device /statistics/system/stats |
# Bulk export of tunnel SLA rows for the last hour from datetime import datetime, timedelta end = datetime.utcnow(); start = end - timedelta(hours=1) fmt = lambda t: t.strftime("%Y-%m-%dT%H:%M:%S") path = f"/data/device/statistics/approutestatsstatistics?startDate={fmt(start)}&endDate={fmt(end)}&count=10000&timeZone=UTC" rows, r = [], m.call("GET", path) while True: rows += r["data"] if not r.get("hasMoreData"): break r = m.call("GET", f"/data/device/statistics/approutestatsstatistics?scrollId={r['scrollId']}&count=10000") # Server-side aggregation: hourly mean latency per tunnel, last 24 h q = {"query": {"condition": "AND", "rules": [ {"field": "entry_time", "type": "date", "value": ["24"], "operator": "last_n_hours"}, {"field": "vdevice_name", "type": "string", "value": ["10.255.1.1"], "operator": "in"}]}, "aggregation": {"field": [{"property": "name", "sequence": 1}], "histogram": {"property": "entry_time", "type": "hour", "interval": 1, "order": "asc"}, "metrics": [{"property": "latency", "type": "avg"}, {"property": "loss_percentage", "type": "avg"}]}} agg = m.call("POST", "/statistics/approute/aggregation", json=q)["data"]
Alarms, events, and webhooks Alarm store
Alarms are correlated conditions with severity; events are the raw notifications behind them. Retrieving alarms over REST means frequent polling, so the Manager can push them instead: a notification rule with a webhook URL delivers an HTTP POST the moment a matching alarm is raised.
| Endpoint | Function |
|---|---|
| GET|POST /alarms | Active alarms; POST body filters by severity, time, site; 26.1 adds site-id, page, pageSize, sortBy |
| /alarms/aggregation, /alarms/severity/summary, /alarms/count, /alarms/topn | Counts by severity (MINOR, MAJOR, MEDIUM, CRITICAL) over time buckets |
| GET /alarms/uuid/{alarm_uuid}, /alarms/notviewed, POST /alarms/markallasviewed?type=active|cleared | Single alarm and viewed state |
| POST /alarms/disabled?eventName=&time= | Suppress an alarm type for 0–72 h (maintenance windows) |
| GET|POST /event, /event/aggregation, /event/severity | Raw events with the same filters |
| GET /auditlog | Who changed what, when |
| POST /notifications/rule, GET /notifications/rules, PUT /notifications/rule?ruleId= | Webhook / e-mail rules; role Settings-write |
| GET /data/device/statistics/alarm/active?startDate&endDate | Bulk export of active alarms for a window |
m.call("POST", "/notifications/rule", json={
"notificationRuleName": "noc-critical-major",
"severity": ["Critical", "Major"],
"alarmName": [], # empty = all alarm types
"devicesAttached": [], # empty = all devices; or [{"system-ip": "…"}]
"webHookEnabled": True, "webhookUrl": "https://hooks.example.com/sdwan",
"webhookUsername": "svc", "webhookPassword": "…",
"emailEnabled": False})
# Backfill: critical alarms in the last 24 h for one site
q = {"query": {"condition": "AND", "rules": [
{"field": "entry_time", "type": "date", "value": ["24"], "operator": "last_n_hours"},
{"field": "severity", "type": "string", "value": ["Critical"], "operator": "in"}]}, "size": 1000}
alarms = m.call("POST", "/alarms?site-id=100", json=q)["data"]Real-time monitoring Reaches the device
Real-time monitoring APIs query device state and information in real time. The Manager opens a NETCONF session to the router over its DTLS (Datagram Transport Layer Security) control connection, runs the query, and returns the result — one device per call. This is the only monitoring path that consumes CPU on the router, and it shares the control channel with configuration pushes and OMP (Overlay Management Protocol) updates.
The scale math: 10 endpoints × 500 routers × once a minute = 5,000 NETCONF sessions per minute — a self-inflicted outage on the Manager.
Deployment model matters. On Cisco SD-WAN Cloud (the Cisco-operated SaaS fabric accessed through an API gateway with an API key, Release 20.15+), real-time and bulk APIs are not exposed. On on-prem or Cisco-hosted control components, where you call your own Manager directly, they are available with the standard limits.
The synced twins
Many real-time endpoints have a synced variant that returns the same shape from the Manager NMS cache: /device/system/status vs /device/system/synced/status, /device/bfd/sessions vs /device/bfd/synced/sessions, /device/control/connections vs /device/control/synced/connections, /hardware/alarms vs /hardware/synced/alarms. If you need one device's view without touching it, use the synced twin.
Endpoint catalog (26.1)
All are GET /dataservice/…?deviceId=<system-ip>. Cisco annotates most as "on vEdge routers only"; the same paths generally work on Catalyst IOS XE edges, but test each one — bridge, PPP, and dot1x are Viptela-OS specific.
Application-aware routing, app logs, ARP
device/app-route/sla-class— SLA classes operating on the routerdevice/app-route/statistics— traffic characteristics per operational data-plane tunneldevice/app/log/flow-count,device/app/log/flows— logged packet flowsdevice/arp— IPv4 ARP table;device/ndv6— IPv6 neighbors
BFD, BGP, OSPF, IP forwarding
device/bfd/history,device/bfd/sessions,device/bfd/synced/sessions,device/bfd/summary,device/bfd/tlocdevice/bgp/neighbors,device/bgp/routes,device/bgp/summarydevice/ospf/database,/databasesummary,/databaseexternal,/interface,/neighbor,/process,/routesdevice/ip/fib(26.1 alsov4fib,v6fib),device/ip/routetable,device/ip/mfiboil,/mfibstats,/mfibsummarydevice/ip/nat/filter,/nat/interface,/nat/interfacestatistics
Control plane, OMP, orchestrator (Validator)
device/control/connections,/synced/connections,/synced/connectionshistory,/localproperties,/synced/localproperties,/statistics,/summary,/waninterface,/synced/waninterface,/affinity/config,/affinity/status,/validdevices,/validvsmartsdevice/omp/peers,/synced/peers,/routes/advertised,/routes/received,/tlocs/advertised,/tlocs/received,/services,/summary,/mcastautodiscoveradvt,/mcastautodiscoverrecv,/mcastroutesadvt,/mcastroutesrecvdevice/orchestrator/connections,/connectionshistory,/localproperties,/summary,/validvedges,/validsmartstunnel/transport/connection— DTLS connection status to the Validator
Interfaces, tunnels, IPsec, QoS and policy
device/interface,device/interface/synced,/arp_stats,/error_stats,/pkt_size,/port_stats,/queue_stats,/statsdevice/tunnel/statistics,device/tunnel/gre-keepalivesdevice/ipsec/inbound,/localsa,/outbound;device/security/informationdevice/policer;device/policy/accesslistassociations,/accesslistcounters,/accesslistnames,/accesslistpolicers,/approutepolicyfilter,/datapolicyfilter,/qosmapinfo,/qosschedulerinfo,/rewriteassociationsdevice/qfp/cpustat,device/qfp/memstat— QFP (quantum flow processor) data-plane load on IOS XE
Flow visibility: DPI, cflowd, CloudExpress
device/dpi/applications,/flows,/summary,/supported-applications(DPI = deep packet inspection)device/cflowd/collector,/flows,/flows-count,/statistics,/template; IOS XE:device/cedgecflowd/app-fwd-cflowd-flows,/app-fwd-cflowd-v6-flowsdevice/cloudx/applications— best interface per CloudExpress application
System, hardware, software, users
device/system/status,device/system/synced/status; 26.1:device/system/info,device/featuresupporthardware/alarms,hardware/synced/alarms,hardware/environment,hardware/synced/environment,hardware/synced/inventory,hardware/thresholddevice/software,device/software/synced;device/reboothistory,device/reboothistory/synced;device/crashlog,device/crashlog/synceddevice/users,device/vpn,device/vrrp,device/ntp/associations,device/ntp/peer
Access: cellular, WLAN, DHCP, dot1x, PPP, bridge, multicast
device/cellular/modem,/network,/profiles,/radio,/sessions,/status,/connection; 26.1:/ursp/routes,/ursp/rulesdevice/wlan/clients,/interfaces,/radios; 26.1:device/wireless/statusdevice/dhcp/interface,device/dhcpv6/interface,device/dhcp/serverdevice/dot1x/clients,device/dot1x/interfaces;device/ppp/interface;device/pppoe/session,/statisticsdevice/bridge/interface,/mac,/tabledevice/igmp/groups,/interface,/statistics,/summary;device/pim/interface,pim/neighbor,device/pim/rp-mapping,pim/statistics;device/multicast/replicator,/rpf,/topology,/tunnel- 26.1:
device/dns/defense/info,device/dns/defense/device-registration; endpoint tracker family under the Real-Time Monitoring - Endpoint Tracker Service tag
Sanctioned implementation: an on-demand runbook
# Triggered by an engineer for ONE device. Never scheduled.
def site_snapshot(m, system_ip):
q = f"?deviceId={system_ip}"
return {
"control": m.call("GET", "/device/control/connections" + q)["data"],
"bfd": m.call("GET", "/device/bfd/sessions" + q)["data"],
"sla": m.call("GET", "/device/app-route/statistics" + q)["data"],
"ifaces": m.call("GET", "/device/interface/error_stats" + q)["data"],
"cpu": m.call("GET", "/device/system/status" + q)["data"],
}Troubleshooting tools Reaches the device
Diagnostic operations that go beyond a single query: log bundles, path traces, utilities. Tagged in the spec as Troubleshooting Tools - Device Connectivity and Troubleshooting Tools - Network Wide Path Insight.
| Endpoint | Function |
|---|---|
| POST /device/tools/admintech | Generate an admin-tech bundle; 26.1 body accepts deviceIP, device-type, exclude-cores, exclude-tech, exclude-logs and a custom-commands array such as show version, show platform |
| GET /device/tools/admintechs, /device/tools/admintech/download/{filename} | List and download bundles |
| GET /troubleshooting/control/{uuid} | Troubleshoot control connections |
| GET /troubleshooting/devicebringup?uuid= | Onboarding diagnostics |
| POST /device/tools/reset/interface/{deviceIP} | Reset an interface |
| POST /device/tools/ping/{deviceIP}, /traceroute/{deviceIP}, /nslookup/{deviceIP} | Connectivity utilities executed on the device |
| /stream/device/nwpi/* | Network-Wide Path Insight: trace/start, trace/stop/{traceId}, traceHistory, exportTrace, importTrace, tasks/* |
| /stream/device/umts/* | Underlay Measurement and Tracing Service sessions |
| /stream/device/log/* | Live log streaming sessions (create, search, renew, download) |
| /stream/device/speed | Speed test session; returns sessionId, startTime, renewalTime |
Errors and resilience
HTTP status codes
| Code | Message | Meaning | Client action |
|---|---|---|---|
| 200 | OK | Success | — |
| 201 | Created | New resource created | — |
| 400 | Bad request | Request was invalid | Fix the body; do not retry as-is |
| 401 | Unauthorized | Authentication missing or incorrect | Re-authenticate, retry once |
| 403 | Forbidden | Understood but not allowed | Check XSRF header and x-roles-required |
| 404 | Not found | Resource not found | Check path / release (deprecated in 26.1?) |
| 429 | Too many requests | Rate limit exceeded | Exponential backoff; no Retry-After is sent |
| 500 | Internal server error | Problem with the server | Retry with backoff; capture body for TAC |
| 503 | Service unavailable | Server unable to complete request | Back off; check Manager cluster health |
Two things the spec does not give you. Login and logout signal success or failure in the response body, not the status code — a failed session login is a 200 with an HTML page. And the OpenAPI document declares 400, 403 and 500 on nearly every operation with no body schema, while 401, 429 and 503 are documented only in prose and carry no rate-limit headers. Your client needs its own backoff and its own error parsing.
When the API returns 200 and nothing
An empty data array with HTTP 200 is not an error to the Manager. Work it top-down:
Failure patterns and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| 200 with HTML body on login | Bad credentials or locked account | Check for <html>; never parse as JSON |
| 403 on POST that works as GET | Missing X-XSRF-TOKEN | Send it on every non-GET |
| 403 with valid token | Role lacks the feature (e.g. Settings-write for webhook rules) | Read the operation's x-roles-required; adjust the usergroup |
| 429 | Exceeded 48/min bulk, 2 concurrent bulk, or 100/s general | Serialize bulk calls; backoff 2 s → 60 s |
Empty data from state tables | Sync blocked or statistics collection disabled | Check /device/blockSync and statistics settings |
scrollId returns nothing | Expired after 10 minutes | Restart the time window |
| Task never completes | Device unreachable or variable error | Read activity[] in the task status |
| Real-time call times out | Device busy or control connection flapping | Use the synced twin; check /device/control/synced/connections |
What changed from 20.18 to 26.1
The Monitoring and Troubleshooting changelog ends with an explicit statement that result-API changes broke backward compatibility. Integrations built against 20.x should be reviewed before an upgrade.
Deprecated
/statistics/bfdand all sub-paths/statistics/cflowdand all sub-paths/statistics/deviceand all sub-paths/statistics/system/statsand all sub-paths/accesstoken/…,/refreshtoken/…,/token/…,/device_authorization/…cloud token helpers/dca/*data-collection endpoints (SD-WAN Services)
Deleted
/statistics/cflowd/applications,…/applications/summary,…/device/applicationsPOST /device/tier/{tierName}→ replaced byPOST /device/tier
New
/advisories/insecure-config/summary,/devices,/devices/{deviceIP};/device/insecure-config/statistics/ppl/ctg,/statistics/ppl/reordering(full query family)/device/featuresupport,/device/wireless/status,/device/cellular/ursp/routes|rules,/device/dns/defense/*,/device/system/info,/device/ip/v4fib|v6fib/statistics/download/{processType}/file/{token}/{fileName}
Changed
site-idfilter across/alarms*,/event*,/notifications/rules,/devicepage,pageSize,sortByacross most/statistics/{table}/health/devicesand/health/devices/overviewresponse reshapedGET /alarms/masteroperation renamed getMasterManagerState → getLeaderManagerStatePOST /device/tools/admintechgainedcustom-commands- JWT authentication introduced in 20.18.1; session auth kept
| Need | Use in 26.1 | Replace |
|---|---|---|
| Tunnel SLA / app-route trending | POST /statistics/approute; bulk approutestatsstatistics | /statistics/bfd |
| Flow / application visibility | /statistics/dpi, /statistics/flowlog | /statistics/cflowd/* |
| Device CPU / memory | /statistics/system (/cpu, /memory); bulk devicesystemstatusstatistics | /statistics/system/stats, /statistics/device |
Guidance checklist
- Never schedule real-time calls. Per-device NETCONF pulls over the control plane belong in an on-demand runbook, scoped to one device or site.
- Poll bulk state for "now." One call per table returns the whole fabric. Five-minute cadence covers NOC dashboards; use
/health/devices/overviewfor the headline. - Export statistics in windows. Bulk
/data/device/statistics/{table}with scrollId paging, two concurrent, 48/min per node; use/statistics/{table}/aggregationwhen the Manager can do the maths. - Push alarms, don't pull them. Register
/notifications/rulewebhooks; use/alarmsonly for backfill. - Build a resilient client. Short-lived JWT with refresh; XSRF on every non-GET; backoff on 429 (no Retry-After); re-auth once on 401; inspect the login body.
- Ramp load gradually and watch the Manager. Start with a low request count and interval, increase one dimension at a time, and watch every Manager node for CPU and memory, API latency and timeouts, 429 and 5xx rates, statistics-query and bulk-export duration, and device reachability. Stop when any of them moves. The rate limit is a ceiling, not a target.
- Least privilege. A read-only monitoring usergroup for collectors;
Settings-writeonly for the account that manages webhook rules. - Target 26.1 endpoint names now. Anything calling
/statistics/bfd,/statistics/cflowd,/statistics/deviceor/statistics/system/statswill break on upgrade. - Know which deployment model you have. On-prem and Cisco-hosted control components expose the full API against your own Manager. Cisco SD-WAN Cloud (the SaaS fabric reached through an API gateway) does not expose real-time or bulk APIs and limits the rest to 5–10 requests/second.
- Know when REST is the wrong tool. Sub-minute telemetry at scale is a job for SD-WAN telemetry data collection or the ThousandEyes Enterprise Agent, not a polling loop.
- Start from the companion kit. A Postman/Bruno collection (40 requests across auth, inventory, bulk state, statistics, alarms, real-time, tasks) and a Python
manager_client.pywith all three auth modes, retry policy and paging helpers ship alongside this guide assdwan-api-field-guide-companion.zip. - Don't reinvent the client. The
catalystwanPython SDK,terraform-provider-sdwan, the Pulumi provider, the DevNet SD-WAN Reporting Tool and the two DevNet MCP (Model Context Protocol) servers already wrap these endpoints; Sastre handles configuration backup and restore.
Built from the Cisco DevNet Catalyst SD-WAN Manager API guides (Releases 20.18 and 26.1), the Cisco Catalyst SD-WAN Getting Started Guide, and the Cisco Control Components and Device Management Guide 26.x. Endpoint names reflect Release 26.1; verify against your Manager's /apidocs before production use.