ArchitectureSecurityWebhooksStripeBubble

Your webhook is not the truth. It is only a notice.

Anyone can send you one, and what you do after it arrives is what separates a protected system from an exposed one.

15 min read Galde Cordeiro Silvestre

On this page
  1. What is a webhook?
  2. What does the event prove?
  3. When the sender is unknown
  4. Recognising the origin is not authenticating the message
  5. What the provider gives you
  6. The check that decides: the signature
  7. Authenticating is not knowing what happened
  8. When you cannot confirm
  9. What does this cost?
  10. What does it cost in Bubble?

Picture this.

What is a webhook?

A webhook is an automatic notice one system sends to another. A payment went through, an invoice came due, a document was signed, a delivery left the warehouse. Something happened over there, and your system hears about it over here, without having to ask every minute.

So the webhook is the call from the scene above. The webhook address is your number, the one that receives the call. And just as you pick a number when you buy a SIM card, you create the webhook address in your system and hand it to the provider, which starts calling it every time something happens. The one placing the call is the provider (or should be), and your system is the one receiving it.

Except that address is not secret. It may sit in the provider's settings, show up in logs and in documentation, and it can leak the way a phone number leaks.

That is why a notice is not a verdict, and almost every webhook problem comes from forgetting that difference: "picking up" and doing whatever the "voice on the other end" asked for. A payment event arrives, the system reads it, marks the subscription as paid, unlocks the account and sends the welcome email. It works, and it keeps working for months, until the morning somebody has access they never paid for and nobody can explain how.

What does the event prove?

In the same way, the webhook gives notice, but picking up is not trusting: the source is what confirms.

The system does what it heard Paymentprovider the payment went through Your system Accountunlocked Nobody checked who was speaking. The system acted on a claim. The system calls back Paymentprovider Your system something happened, maybe calling back: how is this payment now? the answer from the source paidunlock the account refundeddo nothing The answer decides, and one of the two answers is to do nothing.
The same call, two systems. The one below calls back on the official number, and the answer decides what happens: sometimes unlock the account, sometimes do nothing.

Here is a concrete case: a customer pays, changes their mind and asks for a refund within the same minute, which is two calls, two notices. If your server is slow, or the queue backs up, they can land out of order. The system that believes what it received unlocks the account on the second notice, and it stays unlocked even though the payment was refunded. The system that asks the source gets "refunded" both times, and does nothing, which is exactly right.

When the sender is unknown

Ignore the call. The system refuses the event and keeps nothing about it. That is fine when the cost of being wrong is low.

Pick up only to find out who it was, then hang up without doing anything. That has a name in software: it is a record, a log entry. You write down that the event arrived, what address it came from and at what time, and you act on nothing. It looks like very little, but it is what saves the investigation three weeks later, when the question is "had this happened before?".

The third decision is the one that cannot be defended, and it is the one systems pick: doing what the voice asked.

Recognising the origin is not authenticating the message

In software, when it comes to webhooks, it is the same three problems. The source address can be forged, the content of the message can be altered on the way, and the secret that authenticates the call can leak.

This is where the improvised defenses show what they are worth, and also where they stop. They turn up in every language and every tool, because they all come from the same wish to stay away from cryptography, or from not knowing there was anything else. There are five of them, and they almost always show up together.

Creating an address nobody can guess (security through obscurity). A long random name, betting that nobody will ever hit it. But people forget it is probably sitting in several places already: the public documentation, the Swagger page, the browser history, the server log.

Adding a secret to the URL (query string token). The address carries a random parameter, so only whoever knows it can send it, and only whoever knows it can check it on arrival. It works like a password agreed between the two sides and told to nobody else. But remember: it can be stored somewhere not so safe, and it can leak or be intercepted.

Allowing only known addresses (IP allowlist). Some keep a list of known IPs: good as a second layer, poor as the only one. The list changes, the provider's infrastructure moves, and you find out when the processes stop.

Only checking that the event exists (lookup by ID). The system takes the ID from the message and asks the provider whether it is real. It is better than the ones above, because it is the only one that talks to the source. But that answers "this call existed", not "the person speaking is who they claim to be".

Rejecting anything that arrives late (timestamp freshness, the check on the clock). The system reads the timestamp inside the message and throws away whatever is old. It helps against accidental retries and against a stored event being sent again, but the timestamp travels inside the message itself: whoever forges the message writes the timestamp too. It only becomes a real defense when that timestamp sits inside the signature, which is the subject of the next section.

None of these layers is useless, and together they do add something: stacking defenses has a name, defense in depth, which is the name for "not betting everything on one layer". But none of them is enough on its own, and the mistake is using one of them as the lock.

Proves who sent it Proves nothing changed Signature verification recompute it on the raw body, and compare HMAC An obscure endpoint obscurity is not a boundary security through obscurity A token in the URL travels in logs and screenshots query string token IP allowlisting good second layer, lists change IP allowlist The event ID exists confirms it is real, not who sent it lookup by ID Rejecting what arrives late the timestamp travels inside the message timestamp freshness
One check, five habits. Only the first one answers both questions, and that is why it is the one that decides.

What the provider gives you

The five above are what we improvise when we believe there is no choice. But almost every serious provider offers something real, and it pays to know what each one proves before picking.

What the provider offersWhat it proves
A fixed header or Basic auth, configured by youThat whoever called holds the credential. Nothing about the content or the age of the message.
An HMAC signature in a header (Stripe, GitHub, Shopify, Slack, Twilio)Authorship and integrity. With the timestamp inside the signature, the age as well.
An asymmetric signature or a signed JWT (PayPal, Apple, Google)The same, checked with a public key, so not even your side can forge a message.
Encrypted content that you decrypt on your side (Branch)That whoever wrote it held the key, and that nobody along the way read the content.
mTLS, with a certificate on the calling sideThat the connection came from someone holding a certificate you accept.

Two lines in that table get mistaken for the others, and the mistake changes the choice.

The first one is the credential. It answers a single question: "does the caller hold the key?" The signature answers a different one: "is this message exactly this message, and did whoever wrote it hold the secret?" Neither covers the other. If the key shows up in a log and somebody copies it, that person makes up a ten thousand dollar approved payment and sends it with the correct credential, because the credential says nothing about the content behind it. With a signature, that invented content does not add up against the secret and the message is rejected.

The second one is encrypted content. Encrypting scrambles the message so that only whoever holds the key can unscramble it. That stops anyone along the way from reading what is written, and stopping someone from reading is not the same as stopping them from changing. These are two different guarantees: one is secrecy, the other is integrity. In practice they almost always come together, because the modern way of encrypting already flags any alteration, and that has a name, authenticated encryption. But it is the norm, not a guarantee: if the provider only encrypts, check the signature too, whenever it offers one.

And there is a line that is not the provider's, it is yours: if the system always calls back before acting, a forged call costs a useless query instead of a wrong decision. That is called reconciliation, it does not replace the signature, and it shrinks the damage of anything that slips through the checks.

The check that decides: the signature

In software that is signature verification, and the mechanism has a name you will meet in the documentation of every provider that offers it: HMAC.

The provider runs a calculation that mixes the exact message it is sending with a secret key only the two of you know, and sends the result along with it. On your side, the same calculation over the same message, with the same secret key, has to produce the same result. If it does, two facts are settled at once: (1) whoever spoke knows the secret key; (2) the message arrived intact, not a single character of what was written changed along the way. In an asymmetric signature the idea is the same, with only the keeping of the key changing.

So this is a check of a different nature. The defenses above lower the risk, and this is the only one that proves both things, who sent it and that nothing changed, which is why it is the most critical and the most reliable, and stands above all the others.

And there are two main details that decide whether it works:

It checks the text as it arrived. Anything that tries to rewrite or alter the message along the way makes the comparison fail on a message that was perfectly legitimate, and the symptom is always the same: the check rejects something true and nobody can see why. The comparison fails because the message was modified, whether anyone meant to or not.

It has no expiry date. A valid signature stays valid forever.

That opens the door to what is called a replay attack: using a call that was validated in the past to run a process in the future. The defense is the same timestamp freshness check from the earlier list, which here is finally worth something: the timestamp sits inside the signature, and for the signature to stay valid the timestamp cannot be altered, remember? So whoever forges the message has to keep the same content, the same timestamp included. What you choose is the tolerance window, which is how late a webhook is still allowed to arrive. Some providers acknowledge that delays happen and even suggest a value for it in their documentation. Stripe's official libraries, for one, ship with five minutes of tolerance, and its documentation warns against dropping that to zero: zero looks like the strictest setting of all, and it turns the recency check off entirely.

Authenticating is not knowing what happened

Three habits carry the rest.

Confirm at the source. Before anything that involves money or access above all, fetch the object from the provider by its ID and act on what comes back. In software this is called reconciliation. The event says a payment happened; the query confirms the current state of that payment, whether it really went through or was reversed two minutes ago. And there is a second reason, less often remembered: arrival order is not always guaranteed, and sometimes a subscription cancellation arrives before the update that still said active. Whoever writes state from the event lets the last one in win, even though it happened first. So the answer from the source is the only thing that does not depend on order.

Expect the same thing to arrive twice. It is routine, not a defect: the provider resends when your server is slow to answer and the network delivers the same message twice, or simply while somebody is setting the integration up and testing it. This is called idempotency: running the same operation twice leaves the system exactly as it was after the first time. What makes it possible is the event identifier, stored on arrival, also called the idempotency key: the second arrival finds that record and does nothing, which is what stops the card from being charged again. And whenever idempotency kicks in, always answer that it went fine, even having done nothing: refusing a repeat with an error tells the provider the event failed, so it sends it again, which is exactly the opposite of what you want.

Record what happened. It can be at the start and the end of the process, or the whole path. One record (or log) per event, written the moment it arrives and updated at every check, depending on the use case. That record is your audit trail. You should not want to know only whether the process failed, but why it failed, how far it got, and what it believed when it stopped, and a log that only keeps successes does not answer that.

When you cannot confirm

The last principle sounds wrong until the day it saves you. When a check cannot decide, because data is missing or something does not add up, the rest of the flow should do nothing beyond recording the problem for somebody to look at.

Letting the flow carry on would mean acting as if it had worked, as if an unpaid account had gone through. Refusing in silence is worse still, because silence looks exactly like the server being down, and nobody investigates a problem they do not know exists, a system that never complains.

The event that stopped and was recorded has a name too, dead letter queue, which is just a corner where a message that could not be handled waits for a person to review it and decide, instead of vanishing. It sits there with enough context to run again once the cause is fixed.

What does this cost?

None of the security steps listed in the diagram below is a wild invention, and all of them have a name: HMAC verification, timestamp freshness, reconciliation, idempotency, audit trail, fail closed, among others. They are steps that should become habit, and they usually show up whether the integration is payments, deliveries, document signing or anything else that sends you an event and hangs up.

And yes, setting all of this up properly is more work than just taking the call, reading the message, swallowing any error in silence and moving on. But that is one of the main differences between a system that is correct, secure and auditable and one that is quietly ignoring problems, about to meet a vulnerability in the most painful way there is.

A webhook is a notice, not a verdict, and it calls for caution.

What does it cost in Bubble?

In Bubble, in the normal mode, if you decide to use the main security layers that are natively possible, including the optional ones, there will be basically eight of them, since HMAC signature verification does not exist natively in Bubble.

All eight have to be written and maintained by you, and each failure point only reaches your audit trail if there is an extra workflow written to log that point.

The four high ones, HMAC signature verification, timestamp freshness, idempotency and reconciliation, are not alternatives to each other: dropping one does not save work, it reopens the gap that one closed. The rest are reinforcement, and a system can live without them.

Some people, however, use the secret in the URL as the alternative to the HMAC signature, thinking it is 100% safe, when in fact it is extremely vulnerable, above all when it stands alone, and that happens in plenty of cases out of not knowing or out of carelessness with the layers.

Others use event reconciliation to make sure the event exists and to check its current state. That really is good practice, but it does not stop a forged event from triggering the reconciliation for nothing, or from setting off other processes down the chain if there are other verification or validation gaps in the workflow.

With the Webhook Sentinel (Stripe) plugin, five of those checks are done by it and one stops being necessary. On two others it does not decide, but it hands over what they need: the event identifier, which is the key used to check idempotency, and the object identifier, which is used for the reconciliation lookup. The rest stays yours.

On top of that, the recommended records stop being a decision at every step, because the plugin returns the specific reason for the failure, which lets you set up your audit trail at a single point of the workflow, after the verification, instead of one point per check that could fail.

Recommended setup:

Nine checks in a row, all of them written and maintained by you.

The plugin covers five, retires one, and helps with two, handing over the identifier they need. The rest stays yours.

The event arrives nothing has been checked yet Is the signature valid? on the raw body, before parsing HMAC High covered by the plugin Not available natively in Bubble no Does the URL secret match? the signature covers it query string token Low no longer needed no Is it recent? six hours later it is still signed timestamp freshness High covered by the plugin no Is it new? the second arrival does nothing idempotency High the plugin helps no Does the source confirm? act on the answer, not on the event reconciliation High the plugin helps no Is the IP on the list? refuse before spending any work IP allowlist Medium covered by the plugin no Does it carry what I need? the id, the reference, the environment minimum data Optional covered by the plugin no Is it the right environment? test does not touch production environment isolation Optional covered by the plugin no Is it within the limit? flooding, not repeats rate limiting Optional no Continue the workflow unlock, charge, send business logic Log record (recommended) optional at every point, and without it the refusal vanishes five covered, one reason audit trail invalid signaturewrong secretoutside the windowalready seensource disagreesIP refusednot enough datawrong environmentover the limit End the workflow nothing is done fail closed
  • Highnot up for negotiation
  • Mediumreinforcement worth having
  • Optionalyour call
  • Lowbarely counts
The whole flow, with the criticality of each check. Nothing continues without passing the ones before it, and nothing that stops disappears quietly.

Any one of these checks, on its own, can take seconds or minutes to write by hand, but that is not the main point. The main point is that HMAC signature verification does not exist natively, and another one is reuse and maintenance. How many of these checks will still be there six months later, across every webhook the app has, across every app the person manages?