Skip to content
Webb Technologies

Custom software · System integration

Integrating two business systems when one vendor's API is poor: polling, webhooks, files and reconciliation

One system has a modern, documented API. The other has an API that returns fifty records a page with no way to ask what changed since yesterday, a rate limit that runs out mid-morning, and no notice when something is deleted. Somebody writes a script that copies everything every night, and for a while it works. Then a record is missed, another is created twice, and the two systems quietly disagree until a customer or an auditor notices. This guide covers how to design an integration around a poor API so that it moves the right records, recovers from failures on its own, and can prove the two systems agree.

Webb TechnologiesPublished 11 min read

Drafted with AI assistance; facts checked against any sources cited. General guidance, not advice for your situation; verify before relying on it.

On this page

In short

  • List exactly what the weak API can and can't do (change filters, webhooks, limits, deletes, write safety) before choosing a pattern.
  • Decide which system owns each record and field, and how records are matched across systems, before moving any data.
  • Use webhooks where they exist but back them with polling; poll with a watermark and an overlap window so nothing is missed and repeats are harmless.
  • Make every write safe to retry, send what can't be fixed automatically to an exception queue a person works, and never retry validation errors forever.
  • Reconcile on a schedule: compare counts, keys and key fields between the systems, report differences to an owner, and fix only the safe ones automatically.

What a poor API looks like

"The API is bad" isn't specific enough to design around. Write down each gap, because each one has a different workaround.

Common API gaps, what they break, and a workaround

No "changed since" filter

What it breaks
You can't ask for only new or updated records
Workaround
Windowed comparison of recent records, or a scheduled export with modified dates

No webhooks, or unreliable ones

What it breaks
Changes aren't pushed; missed events are lost
Workaround
Polling, or webhooks with a polling backstop

Tight rate limits

What it breaks
Large syncs fail partway; other integrations get starved
Workaround
A request budget per integration, backoff, and syncing only what changed

Deletes aren't visible

What it breaks
Deleted records live on in the other system
Workaround
Periodic comparison of the full list of IDs

Unstable paging

What it breaks
Records shift between pages while you read; some are skipped or repeated
Workaround
Sort by a stable key, or page by ID ranges

Writes aren't idempotent

What it breaks
A retried create makes a duplicate
Workaround
Store your own ID on the record and look it up before creating

Fields or objects not exposed

What it breaks
The data you need isn't reachable at all
Workaround
Ask the vendor; a scheduled export or reporting database may have it

Unannounced changes

What it breaks
A field is renamed or a format changes and the sync breaks
Workaround
Validate every response against an expected shape and alert on mismatches

Before you write any code

  • Ask the vendor specific questions. Is there a change feed, webhook or export you haven't found? What are the rate limits for your license tier? Is there a sandbox? How are API changes announced? Some vendors tie API access or higher limits to particular editions or add-ons, so confirm what your contract includes, as of the date you ask.
  • Check the terms. Reading a vendor's database directly, or automating its screens, may breach your agreement or void support.
  • Decide who owns what. For each record type and field, name the system of record. Two systems that both accept edits to the same field will disagree, whatever the integration does.
  • Agree how records match. A shared identifier, or a cross-reference table the integration keeps, links a record in one system to its partner in the other. Matching on names or emails alone creates duplicates.
  • Size the need. How many records change per day, and how soon does the other system need to know: seconds, minutes or by the morning? That answer rules patterns in or out.

Choosing a pattern

Integration patterns for a weak API (combine them where it helps)

Webhooks with a polling backstop

Fits when
The vendor sends events, but you can't rely on every one arriving
What to watch
Signature checks, duplicates, events out of order, and the backstop catching what was missed

Polling with a watermark

Fits when
The API can filter or sort by a modified time or sequence number
What to watch
Clock differences, records changed in the same second, and rate limits

Windowed comparison

Fits when
No change filter, but you can list recent records cheaply
What to watch
The window must be wider than any delay; deletes still need a separate check

Scheduled file export

Fits when
The vendor offers exports more complete than its API
What to watch
File completeness, formats that change, and data that is only as fresh as the last export

Direct database read

Fits when
Only where the vendor permits it, for reading
What to watch
Support terms, schema changes on upgrade, and load on the vendor's database

Screen automation, where a script drives the vendor's web pages, is a last resort: it breaks when a page changes and may conflict with the vendor's terms. If it's the only route, treat it as temporary and keep pressing the vendor for a supported one.

A small integration service in the middle

integration / weak API to target systemdiagram
Changes leave the source system with the weak API through webhooks where it sends them, polling with a watermark, or a scheduled file export. A small integration service receives or fetches them, stores each raw record or event as received, and checks it against the expected shape. It maps fields using the cross-reference of IDs, then writes to the target system with a key that makes retries safe. Temporary errors are retried with backoff; anything that can't be fixed automatically goes to an exception queue with a readable reason for a person to resolve and replay. A scheduled reconciliation job compares the two systems and reports differences, and every run, retry and exception is logged and monitored with alerts to the integration's owner.

Each system stays the record for its own data. The service moves and checks records; it doesn't become a third copy people edit.

  • Store what you received. Keeping each raw record or event lets you replay after a mapping fix and answer "what did the vendor send us?" without guessing.
  • Keep mapping rules in one place. Field mappings, code translations and defaults live in the service, versioned and tested, not scattered across scripts.
  • Small and boring beats clever. A few functions, a queue and a table of IDs, defined as code and deployed automatically, are easier to hand over than a large framework. Our guide to AWS foundations for your first custom apps covers the accounts, environments and monitoring around a service like this.

Polling that doesn't miss or repeat records

  • Keep a watermark. Record the latest modified time (or sequence number) successfully processed, and ask for changes after it on the next run. Advance it only after the batch is written.
  • Overlap the window. Ask for changes from a little before the watermark, to catch records saved in the same second or delayed by the vendor's own processing. Repeats are then expected, so writes must be safe to repeat.
  • Sort by time, then ID. A stable order means a record isn't skipped when several share a timestamp across a page boundary.
  • Watch the clocks. Use the vendor's timestamps, not your server's, and know which time zone they're in.
  • Budget the rate limit. Decide how many calls per hour this integration may use, leave room for other integrations and people, slow down when the vendor says to (for example a Retry-After header), and back off with some randomness after errors.
  • Catch deletes separately. If deletes aren't visible, compare the full list of IDs on a slower schedule and handle the missing ones by your agreed rule: delete, archive or flag.

Webhooks you can trust

  • Verify the sender. Check the vendor's signature or shared secret on every call, and reject anything that fails.
  • Answer quickly, work later. Store the event and respond at once; do the processing from a queue. Slow responses can make the vendor retry or disable the webhook.
  • Expect duplicates and disorder. Deduplicate by event ID, and don't assume events arrive in the order things happened. Where it matters, fetch the record's current state from the API rather than trusting the event body.
  • Back it with polling. A light poll or comparison on a schedule catches events that never arrived, such as those sent while your endpoint was down.
  • Watch for silence. No events for longer than expected can mean the vendor disabled the webhook. Alert on that as you would on errors.

If the weak side only offers files, the same principles apply to each file: confirm it's complete, check its shape, keep the original and process it once.

Retries, idempotency and partial failure

What to do with each kind of failure (agree the rules with the system owners)

Temporary

Example
Timeout, rate limit reached, vendor briefly down
What the service does
Retry with backoff, up to a limit, then send to the exception queue

Data problem

Example
Required field missing, code not in the mapping, record fails validation
What the service does
Don't retry; send to the exception queue with the reason, for a person to fix at the source or in the mapping

Conflict

Example
The target record was changed by someone since the last sync
What the service does
Apply the ownership rule for that field, or send to the exception queue if the rule doesn't decide it

Unknown outcome

Example
The write timed out and you can't tell whether it happened
What the service does
Look up the record by your stored ID before trying again
  • Make writes safe to repeat. Use the target's idempotency key if it offers one; otherwise store your own ID on the target record and look it up before creating.
  • One record's failure isn't the batch's. A bad record goes to the exception queue; the rest of the batch continues.
  • Exceptions a person can act on. Each entry shows the record, the reason in plain language, and a button to replay once fixed. Age and count of open exceptions are the numbers to watch.

Reconciliation: proving the two systems agree

A sync that reports no errors can still have missed records. Reconciliation compares the two systems directly, on a schedule, and turns silent drift into a list someone can work.

Reconciliation checks (adjust to your records and volumes)

Counts

How
Records of each type in each system, for a period or status
Who sees the result
The integration's owner, on a status page

Keys

How
IDs present in one system and missing from the other
Who sees the result
The owner, as a list with links to both records

Key fields

How
A comparison of the fields that matter most, such as status, amounts or dates
Who sees the result
The business owner of those records

Exceptions

How
Open items, their age and reasons
Who sees the result
The owner and anyone who fixes source data
  • Fix the safe ones automatically. A record missing from the target because an event was lost can be re-sent. Differences in fields owned by people go to a person.
  • Track drift over time. A count of differences found per run shows whether the integration is getting better or worse, and whether a vendor change broke something.

Monitoring and ownership

  • Alert on the absence of success, not only on errors: no successful run in an agreed period means something is wrong even if nothing failed loudly.
  • Alert on exception age and rate-limit headroom, so problems are seen before the business notices.
  • Watch for vendor change notices. Someone subscribes to the vendor's API change announcements and deprecation dates, and the shape checks catch what wasn't announced.
  • A short runbook. How to pause the sync, replay exceptions, rerun reconciliation and rotate credentials, written for whoever is on call.
  • Named owners. A technical owner for the service and a business owner for each record type, who decides what happens to differences.

Keep the integration's credentials in a secrets store, scoped to only the objects it needs, and rotate them on a schedule. Our guide to SSO and least-privilege access for internal apps covers service identities and scoping.

An integration platform, a vendor connector or a small custom service

Routes to the integration (check current features and pricing with each vendor)

A connector from either vendor

Where it fits
One vendor offers a supported connector to the other system
What to check
Which fields and objects it covers, how it handles failures and deletes, and whether you can see its logs

An integration platform you already license

Where it fits
Your company runs one and has people who maintain it
What to check
Whether it handles watermarks, idempotent writes and reconciliation, or only simple flows, and its per-run or per-connection pricing

A small custom integration service

Where it fits
The weak API needs careful polling, retries and reconciliation that a connector doesn't offer
What to check
Ownership of the code, hosting, monitoring, support after launch, and who maintains the mappings

Whichever route you choose, the groundwork carries over: the list of API gaps, ownership of each field, the ID matching rule, the failure rules and the reconciliation checks. Our guide to custom web apps vs. low-code covers the wider build-or-buy decision.

Two systems that should agree but don't?

Bring the vendor's API documentation, a few examples of records that went wrong and a note of how often each system changes. On a scoping call we'll talk through the pattern, retries, reconciliation and what a fixed-price first phase would cover.

Request a scoping call

Not ready to talk yet? Draft a project brief first 

Questions to settle before you build

Take this to the system owners and IT

  • What exactly can the weak API do and not do, confirmed with the vendor for your license?
  • Which system owns each record type and field?
  • How are records matched between the systems, and where is the ID cross-reference kept?
  • How many records change per day, and how soon must the other system know?
  • Which pattern, or combination, fits: webhooks, polling, windowed comparison or file exports?
  • Which failures are retried, which go to a person, and who works the exception queue?
  • Which reconciliation checks run, how often, and who acts on the differences?
  • Who owns the service, the mappings and the vendor relationship after launch?

Our experience includes RESTful and GraphQL APIs, serverless architectures on AWS, database design on SQL Server, PostgreSQL and MySQL, OAuth and OIDC authentication, and monitoring and alerting with CloudWatch. See custom software development for what we build, and how we work for the fixed-price process.

Sources

Want to talk through your version of this?

A 30-minute call. You leave with a clear approach and the real risks, whether or not you hire us.