Docs

Docs

Monitor an endpoint

Run an endpoint connection's check on a schedule, from our cloud or from your agent, keep the results for 90 days, and get an event when the target goes down and when it recovers.

Pre-release

Monitors are new (RFC 0037). The routes and fields below are the contract; the ceilings may rise.

What a monitor is

A monitor runs a connection's check action every N seconds and keeps each result. A check is one GET that passes when the answer has the expected status (any 2xx unless you name one) within the time limit you set. From the runs the platform works out the monitor's health, its uptime and its latency, and it sends an event when the health changes.

A monitor belongs to its connection. It inherits the connection's account, its executor and its credential, and it is deleted with the connection. Only kinds with a check action can be monitored; today that is endpoint.

Where the check runs

The executor is the connection's, fixed when the connection is made:

  • Our cloud. For a public HTTPS endpoint inside a domain your account has verified. The platform resolves the name, refuses private addresses and checks again that the domain is still verified before every run.
  • Your agent. For anything the agent can reach: a host inside your network, a staging name, a port the internet never sees. The agent's own policy decides, and it reports the outcome, never the body. See Installing an agent.

Make one

In the console. Open Monitors (Settings › Reliability) and choose New monitor, or open an endpoint connection and choose Monitor this. Pick how often (every 1, 5 or 15 minutes), the expected status and time limit if any, and how many failures in a row mean down.

With the API. A key or token with connections:write:

iohr api POST /v1/accounts/orgs/<account>/connections/<connection>/monitors \
  -F interval_secs=60 \
  -F 'params={"path":"/healthz","expect_status":200,"max_ms":2000}' \
  -F fail_after=2

The answer is the monitor, with its id (mon_…). The first run starts within a minute. name is optional; without it the monitor shows the connection's name.

FieldMeaning
interval_secsseconds between runs, 60 at least
paramspath (after the connection's URL, empty for the URL itself), expect_status (100 to 599, any 2xx when left out), max_ms (up to 30000)
fail_afterfailures in a row before the monitor is down, 1 to 5, 2 when left out
nameoptional, a label for people

Health

A monitor is up, down or unknown:

  • It goes down after fail_after failed runs in a row, and back up at the first run that passes. One slow answer does not page anyone unless you set fail_after to 1.
  • It is unknown before its first run, and while the agent that runs it is offline. The target was not seen, which is not the same as the target failing, so an offline agent never makes a monitor down.
  • The platform pauses a monitor whose setup is wrong instead of failing it forever: the domain is no longer verified, the connection is paused or deleted, or the agent was revoked. paused_reason says which. Fix the cause and set state back to active.

Read it

RouteScopeWhat it does
POST /v1/accounts/orgs/{org_id}/connections/{connection_id}/monitorsconnections:writemakes a monitor
GET /v1/accounts/orgs/{org_id}/monitorsconnections:readthe account's monitors, paged; filter with connection_id
GET /v1/accounts/orgs/{org_id}/monitors/{monitor_id}connections:readone monitor
PATCH /v1/accounts/orgs/{org_id}/monitors/{monitor_id}connections:writechanges name, interval_secs, params, fail_after, state
DELETE /v1/accounts/orgs/{org_id}/monitors/{monitor_id}connections:writedeletes it and its runs
GET /v1/accounts/orgs/{org_id}/monitors/{monitor_id}/runsconnections:readruns, newest first, paged; from an RFC 3339 time
GET /v1/accounts/orgs/{org_id}/monitors/{monitor_id}/summaryconnections:readhealth, uptime over 24 h, 7 d and 30 d, p50 and p95 latency

A run is {run_id, at, status, ok, latency_ms, status_code, error_class}; status_code and error_class are left out when there is none. A run never holds a request or response body. Creating and changing a monitor honour Idempotency-Key.

iohr api GET /v1/accounts/orgs/<account>/monitors/<monitor>/summary

Get told

A change of health is an event:

  • monitor.down: the monitor went down.
  • monitor.recovered: it passed again after being down.

Both carry ids only (account_id, monitor_id, connection_id); read the summary or the runs for the rest. They are delivered like every other event: subscribe a webhook endpoint to them, or follow them over MQTT or the event stream.

iohr api POST /v1/webhooks/endpoints \
  -f url=https://example.com/hooks/uptime \
  -F 'event_types=["monitor.down","monitor.recovered"]'

Limits

LimitValue
Shortest interval60 seconds
Monitors per account20
Runs kept90 days
Time limit of a check30 seconds

Runs started by the scheduler are not charged in units yet; the API calls that make, change and read monitors are, like any other call.