This benchmark report describes a fixed synthetic test, not a production-agency deployment or a promise about how this stack will perform for every agency.

I tested an agency client-onboarding workflow built with Make, HubSpot Free, Gmail, and Google Calendar. For each valid benchmark input, the workflow was expected to create the CRM state and four onboarding tasks, produce a welcome email and internal handoff, and create a kickoff calendar artifact.

The short version: the normal valid inputs worked consistently in the measured series. The separate edge tests, however, exposed weak input validation and idempotency behavior. That distinction is the important result.

What the 20-run record actually contains

The benchmark retained 20 main_initial records. They are not the same thing as 20 valid measured runs.

  • Run 1 was an operator-invalid record: the Make scenario was inactive when the fixture was submitted. The webhook accepted and queued the payload, but the workflow did not execute; the queued item was inspected and deleted.
  • The remaining 19 records were valid main runs.
  • All 19 valid main runs succeeded: 19/19 (100.00%).

For those 19 valid runs, Make’s vendor-reported scenario duration had a median of 5 seconds, a minimum of 4 seconds, and a maximum of 5 seconds. The timing runs from webhook-triggered execution start to scenario completion; post-run evidence review was excluded.

That 19/19 observed pass rate applies only to this fixed synthetic benchmark set. It is not a population reliability estimate, a guarantee of production performance, or evidence that every agency intake will behave the same way.

For more on what this site treats as a benchmark measurement, see the methodology.

The edge tests changed the conclusion

The main series used valid inputs. Separate edge tests intentionally examined how the workflow behaved when required data was missing, identity conflicted, or a submission was repeated. They are not included in the 19/19 main-series pass rate.

Edge test Observed outcome Why it matters
Missing required email Propagated The input created partial HubSpot state before Make failed at Gmail because the recipient was missing.
Duplicate client identity Duplicated Two submissions with the same client identity created duplicate company and downstream artifacts rather than being safely rejected or updated.
Malformed website Propagated The literal value not a URL was accepted and passed through the workflow.
Exact repeated submission Duplicated Two identical submissions created a second set of downstream artifacts rather than behaving idempotently.

This is why “19/19” is not the whole story. The workflow handled normal valid inputs consistently in the measured series, but it did not reliably stop bad input before partial propagation or prevent duplicate work after identity and replay cases.

The retained observations recorded active cleanup time of 1.25 minutes for the missing-email case, 2.17 minutes for duplicate client identity, 1 minute for the malformed website, and 1.67 minutes for an exact repeated submission. Those are recovery observations for these specific test cases, not a general maintenance-time estimate.

The exact-repeat finding deserves a narrower evidence note: its downstream duplicate counts are supported by a retained structured observation. No retained vendor-side downstream screenshot exists for that edge case.

What I would change before relying on this workflow

The next implementation should add validation before any CRM, email, or calendar action can run. At minimum, it should reject or quarantine an intake without a required email and validate a website field before persisting it.

It should also define an idempotency strategy. The duplicate-identity and exact-repeat tests show that a valid-looking intake can still create duplicate company records, tasks, messages, and calendar artifacts when it is submitted again. A stable external key, a lookup-before-create step, or another explicit deduplication design would need its own test rather than being assumed.

A future comparison, HighLevel vs Make + HubSpot for Agency Client Onboarding: Same Workflow, Same Test, should use the same external task and make the same input-validation and idempotency behavior visible.

Cost and account context

The retained account evidence supports a capture-date minimum recurring software cost of $0/month, with a 2026-09-21 basis. That is not a reconstructed historical price or a claim about a business-grade deployment.

  • Make was on the Free plan: $0/month and 1,000 credits/month.
  • HubSpot’s current subscription was Free.
  • The workflow used consumer Gmail and Google Calendar; the supported incremental recurring software cost for those components was $0. No Google Workspace subscription is asserted.

Make showed 446/1,000 credits used in the current cycle at capture. The operator states that cycle’s Make account use was only for AWB/affiliate experimentation, so those 446 credits may describe current-cycle AWB experimentation usage. They are not a measurement of the 20-run main series alone.

metered_test_cost_usd is unavailable: there is no retained billing invoice, usage-charge record, paid-overage record, or credit-purchase record that establishes historical metered cash charges. setup_active_minutes and setup_wait_minutes are also unavailable / not measured; they were not reconstructed.

Affiliate disclosure: the Make link below is an affiliate link; the author may receive a commission if a reader purchases through it, at no additional cost to the reader. The HubSpot link is a normal, non-affiliate link. Affiliate relationships do not determine the reported measurements or conclusions. See the site’s affiliate disclosure.

Evidence and publication limits

Representative vendor-side evidence was retained for a successful main path, including Make execution history, HubSpot company/contact and task state, Gmail messages, and a Google Calendar kickoff. It was also retained for the missing-email, duplicate-client, and malformed-website edge behavior. The site currently has no established public article-image pipeline, so this article does not copy account, billing, or vendor-console screenshots into public assets.

The completion review marks qualified publication readiness as passed, but strict frozen benchmark completion was not achieved because setup-time measurements were not retained. This article therefore does not claim general reliability, ROI, labor savings, or cost savings. It reports the retained test record and the specific weaknesses the edge tests exposed.