Delivery retries and dead-letter queues

Turn on the queue and a request that cannot be delivered waits, is retried and, if it never gets through, lands where you can see it.

Every outgoing REST, SOAP, HL7 FHIR and HL7 MLLP connection has a Delivery tab in the Dashboard. What you set there decides what happens to a request that could not be delivered:

  • It's retried right away
  • It waits in the connection's queue until it goes through
  • It lands in the dead-letter queue (DLQ), where you can look at it, retry it or discard it

What happens to a failed request

You can configure retries in 3 ways:

  1. Retries - the call is retried a few times, with a growing wait in between, and your service waits for the outcome
  2. The queue - the call is made once and, if it failed, the request is put in the connection's queue and your service moves on, the queue then retries it in rounds until it goes through
  3. The dead-letter queue - a request that a round could not deliver is set aside in the DLQ, from where it's retried later, forwarded to a topic, discarded or kept for you

By default, nothing is turned on, a failed call is reported to your service and nothing is retried.

Retries

A request is retried when:

  • The remote system is unreachable
  • It did not answer in time
  • It answered with HTTP 429
  • No HL7 MLLP acknowledgment came back

Everything else, e.g. an HTTP 500 or a negative HL7 acknowledgment, is a response and reaches your service as it arrived.

In the Delivery tab, click the On failure line to set the retries:

SettingDefaultWhat it does
Max. retries0How many times a failed request is retried
Wait before the first retry2 secondsEach retry after the first one waits longer, see the multiplier below
Wait multiplier2Each retry waits this many times longer than the previous one
Wait in total at most60 secondsA cap on the total time spent waiting, once reached, no more retries happen

No single wait of this schedule is longer than 8 seconds.

Retry options on a connection

For instance, with 5 retries and a 2-second first wait, a request that keeps failing is attempted six times in about 30 seconds:

Attempt123456
Wait before-2s4s8s8s8s

Rate-limited responses

An endpoint that answers with HTTP 429 may say how long to wait in a Retry-After header, as a number of seconds or as an HTTP date, and the connection waits that long instead of following its own schedule, the 8-second cap does not apply then.

A wait that does not fit in what is left of the total cap ends the retries and the 429 response is returned to your service.

Retry-After is honoured by the retries above, not by the rounds of the queue below.

The queue

Turn on Use queue and a request that could not be delivered is put in the connection's queue and your service moves on.

  • The call is made once, without retries, and any failure puts the request in the queue, an error response such as an HTTP 500 included
  • If the queue already has requests in it, a new one goes straight to the queue, behind them
  • Requests that only read, e.g. a GET, are never queued
  • The queue survives server restarts
  • Each connection has a queue of its own and requests are delivered in the order they arrived
  • Each delivery is a round of attempts under the retry settings above, and a request stays in the queue until a round succeeds or the request expires

A request that a round could not deliver moves to the dead-letter queue. With the DLQ off, it stays at the head of the queue and is tried again, round after round, while everything behind it waits.

The dead-letter queue

Use DLQ is on once the queue is on. A request that a round could not deliver moves to the connection's dead-letter queue with a note of why, when, and how many attempts it took, and the queue behind it carries on.

Action says what happens to each request in the DLQ after Every has passed since it arrived:

ActionWhat it does
Keep in DLQThe default, the request waits for you to retry, forward or discard it in the Dashboard
RetryThe request goes back to the queue for another round, at most Max. retries times, then stays in the DLQ
Forward to topicThe request is published to a pub/sub topic of your choice, Keep the DLQ header says whether the note goes with it
DiscardThe request is deleted

Every defaults to a minute and Max. retries to 3.

Watching the queue and the DLQ

An outgoing connection's row in the Dashboard has a Delivery queue link, which opens a page with two tabs, Delivery queue and Dead-letter queue. Both list the requests in them and when they arrived, and the DLQ tab also shows how many rounds there were and what the last error was. For each request you can:

  • See its body and headers
  • Download it
  • Edit the body before it's retried
  • Discard it, or retry it if it's in the DLQ

From your services

Nothing changes in how a service calls a REST, SOAP, HL7 MLLP or HL7 FHIR connection, but with the queue on, a call returns a result instead of a response. is_ok means the request was delivered and the endpoint's answer is in response, is_in_queue means it's in the queue under msg_id, and error says why if it's neither.

# Get a REST connection ..
conn = self.rest['Billing API']

# .. send the invoice ..
result = conn.post(self.cid, invoice)

# .. and check what happened to it.
if result.is_ok:
    self.logger.info(f'Delivered -> {result.response.data}')

elif result.is_in_queue:
    self.logger.info(f'Queued -> {result.msg_id}')
# Get a SOAP connection ..
conn = self.soap['Immunization Registry']

# .. invoke an operation ..
result = conn.invoke('submitSingleMessage', request)

# .. and check if it's in the queue.
if result.is_in_queue:
    self.logger.info(f'Queued -> {result.msg_id}')

A REST outgoing connection can also take a request straight to the queue, the call returns as soon as the request is stored:

# Get a REST connection ..
conn = self.rest['Billing API']

# .. and queue the invoice.
result = conn.publish(invoice)

With the queue on, a send that got a negative acknowledgment is queued rather than returned, and response is the acknowledgment of a delivered message:

# Get an MLLP connection ..
conn = self.mllp['EHR Main']

# .. send the message ..
result = conn.send(message)

# .. and check what happened to it.
if result.is_ok:
    self.logger.info(f'Acknowledged -> {result.response.ack_code}')

elif result.is_in_queue:
    self.logger.info(f'Queued -> {result.msg_id}')

With the queue on, saving a resource returns the result instead of the saved resource:

# Get a FHIR connection ..
client = self.fhir['FHIR.Sample']

# .. build a patient ..
patient = client.resource('Patient')
patient.birthDate = '1974-12-25'

# .. save it ..
result = patient.save()

# .. and check if it's in the queue.
if result.is_in_queue:
    self.logger.info(f'Queued -> {result.msg_id}')

A FHIR connection can also take a resource straight to the queue, the call returns as soon as the resource is stored:

# Get a FHIR connection ..
client = self.fhir['FHIR.Sample']

# .. the resource to queue ..
patient = {
    'resourceType': 'Patient',
    'birthDate': '1974-12-25',
}

# .. and queue it.
result = client.publish(patient)

See also

PageWhat it covers
REST outgoing connectionsCreating connections, pools, timeouts and OpenAPI import
Error handlingWhat your service receives when an endpoint fails
Enmasse referenceThe same settings as YAML keys, under each connection type
Pub/sub topics and message queuesThe topics a dead-lettered request can be forwarded to

Learn more