Work / Distributed System

Event-Driven Integration Platform

Partners POST into a NestJS API. I check the body, save the event in MongoDB, publish to RabbitMQ, and return. The partner call runs on a worker. Vue only shows whether it landed.

  • Event queues
  • REST APIs
  • Async workers
  • RabbitMQ
  • Retries
  • Dead-letter queue
  • Webhooks
  • Rate limiting
  • Partner APIs

01

Project overview

Other systems post work to this API. I store it, put it on RabbitMQ, and let a worker call the partner. The browser, or a mobile app someone else owns, only waits for the “got it” response.

02

While someone is using it

A partner call inside the user’s request makes them wait, and the call is gone if the process dies halfway. Partners also send the same webhook twice. Without a record of the first one, you get two emails or two orders.

A partner posts an event and gets an accepted response before the downstream call runs. An operator watches that event move from waiting to running to finished, or to dead-lettered. The status screen only reads. It does not send the work again.

  1. A webhook arrives

    The guard checks the signature and the timestamp. The document is saved with an idempotency key, then published. The partner gets “accepted” while the other system has not been called yet.

  2. The worker runs

    It reads the key first. If that event is already finished, it acks the message and skips the email or the order. If not, it calls the partner and writes the result back on the same document.

  3. The same POST again

    The second delivery is a skip. The partner still gets the accepted response. One side effect, not two.

  4. The partner is down

    The worker retries with a delay, then a ceiling. After the last try the message sits on the dead-letter queue and the document stores the last error. The status screen shows failed. An operator can replay it. The screen cannot.

  5. Someone is watching

    Vue reads the document. The states it can show are waiting, running, failed, and dead-lettered, plus the attempt count. Nothing on that page edits the payload.

03

Business problem

Calling a partner API inside the user request makes the user wait, and the call is gone if the process dies halfway. Partners also send the same webhook twice. If the handler is not careful, you get two orders or two emails.

04

Requirements

  • The API accepts an event and responds without waiting for the downstream system.
  • A failed delivery is retried with a limit, then parked where a person can inspect it.
  • The same event delivered twice does not produce two side effects.
  • Partner webhooks are authenticated and rate limited.
  • Operators can see whether a workflow is waiting, running, failed or dead-lettered.

05

Who can do what

A partner can hand work in. An operator can see it fail and send it again. The status screen can read the state. None of them can rewrite an event that was already stored.

Action Partner Operator Status screen
POST a signed webhook Yes No No
Get a response before the partner call finishes Yes No No
See waiting, running, failed, or dead-lettered Own events Yes Yes
Read the last error and the attempt count Own events Yes Yes
Replay a dead-lettered message No Yes No
Edit the stored payload No No No

A second POST with the same idempotency key does not run the side effect again. The partner still gets the accepted response. The worker skips it.

06

How it’s wired

Vue is the screen I use to see status. It reads a NestJS API. The API checks the payload, stores it in MongoDB, and publishes to RabbitMQ. Workers call the other systems and write the result back. Redis keeps rate limits and short locks. After the last retry, the message sits on a dead-letter queue.

  1. Vue.js
  2. NestJS API
  3. RabbitMQ
  4. Workers
  5. MongoDB / external APIs

07

What I used

These are the tools on this project. I didn’t add anything just to fill a stack diagram.

  • Node.js
  • NestJS
  • MongoDB
  • Vue.js
  • RabbitMQ
  • Redis
  • Docker

08

Technical challenges

  • Publisher confirms matter. Writing “queued” in MongoDB before the broker accepts the message creates a lie. Publishing first and crashing before the database write loses the audit trail.
  • Retries need a delay and a ceiling. Immediate requeue loops burn workers when a partner is down.
  • External APIs time out. The worker has to decide whether the remote side actually performed the action before it tries again.

09

Solution

Treat intake as a recorded fact. The API stores the event with an idempotency key, publishes to RabbitMQ, and only then marks the record as queued. Consumers are idempotent: before a side effect, the worker checks whether that key already reached a terminal state. Failures retry with backoff. After the last attempt the message is dead-lettered and the record is marked failed, with the last error attached. Webhooks are verified with a shared signature and rejected when the timestamp is outside a short window.

10

Implementation

NestJS modules split intake, publishing, consumption and partner adapters. Controllers do not call third parties directly. Adapters hide partner-specific payloads so a vendor change does not spread through the workflow. MongoDB stores the event document, attempt history and the idempotency key on a unique index. RabbitMQ uses a work queue plus a dead-letter exchange. Redis limits how many calls a partner receives per window. Vue reads the event document and shows state, attempts and the last error.

11

Testing

Contract tests cover the public intake schema. Integration tests publish a message and assert one side effect when the same key is delivered twice. A downed partner is simulated to assert retry counts and a final dead-letter. Signature tests reject missing, stale and altered webhooks.

12

Deployment

API and worker processes ship as separate containers from one image. RabbitMQ, Redis and MongoDB are external services. The worker is scaled by consumer count, not by adding webhook endpoints. Health checks cover the process, and a dead-letter depth alarm is the operational signal that a partner is stuck.

13

Outcome

The HTTP call stays short. Retries, the idempotency key, and the dead-letter queue are what keep a bad partner from losing the work or doing it twice.

14

What I decided

  • RabbitMQ rather than an in-process queue, because acknowledgements, retries and a dead-letter queue are the product, not an extra.
  • MongoDB for the event log, because each partner payload is shaped differently and the document already matches the attempt history we need to show operators. Relational integrity is not the hard part of this system. Duplicate delivery is.
  • NestJS for the API boundary so modules, guards and queues have a place to live as more partners are added, instead of a single Node file that knows every vendor.

Start a project