Hearing that all 19 runs succeeded may make an automation seem ready to rely on. But for AWB-EXP-001 Implementation A, 19/19 means successful valid runs / valid main runs observed in the main test. It is not an estimate of the probability of success across real-world operation, and it does not mean “100% reliable.”

The subject is customer onboarding built with Make, HubSpot Free, Gmail, and Google Calendar. The test used fixed synthetic data; it did not measure a live agency’s production operation. The configuration and main-test results are covered in the first experiment report. This article reads the frozen records to clarify what can—and cannot—be concluded from that success count.

Start with the denominator behind 19/19

There are 20 retained records from the initial main test (main_initial). Of those, Run 1 is retained as an operator-invalid, non-counting run because the Make scenario was disabled at submission. The webhook accepted and queued the input, but the workflow did not run. The queued item was removed after verification.

Run 1 is recorded as valid_run=false and failure_category=operator_error, so it is excluded from the valid main-test total. It was neither deleted from the record nor counted as a success. That leaves 19 valid main runs, all 19 of which succeeded.

The success determination also did not rely only on Make showing completion. Under the frozen definition, it checked CRM state, four onboarding tasks, the welcome email, a kickoff event, the internal handoff, and execution logs, and required that the necessary outputs contain no unintended duplicates. However, final delivery of email to an inbox was not a main-test completion criterion; the criterion was acceptance by the configured Gmail action.

What this establishes is that the defined success criteria were met in the valid runs of this fixed test set. It cannot be directly reinterpreted as a population reliability estimate covering other inputs or operating conditions.

What happened in boundary cases outside the main test

Boundary cases are separate from the 19/19 total. They cover missing inputs, invalid values, duplicate client IDs, and repeated submissions. In addition to whether valid input can be processed, these records examine what may be created when input is problematic or resubmitted.

missing-email: partial CRM data remained when processing stopped

An input missing the required email address was not rejected or isolated at entry; it propagated to HubSpot. A company, a contact with an empty email address, and four tasks were created before Gmail raised a BundleValidationError because the recipient to was missing. The later calendar processing and internal notification did not run.

The observed outcome is propagated. Stopping on an error is different from preventing incomplete data before it is created. For this case, the record documents manual deletion of the HubSpot data that had already been created.

duplicate-client: changing the company name with the same client ID created duplicates

The valid observation is edge-duplicate-client-20260921-r3-fix. A normal input was first submitted, followed by one with only company_name changed while retaining the same client_id. This was not an identical resubmission; it tested handling of different inputs with the same client ID.

Both executions succeeded in Make, but company records and the subsequent tasks, emails, and kickoff events were duplicated. The contact is shared. The observed outcome is duplicated, not safe rejection or a duplicate-free update.

The frozen configuration document describes an intended design in which a match on an existing awb_client_id stops subsequent creation. Even so, this valid boundary test observed duplication. Design intent cannot be treated as confirmed behavior. Nor does this record alone establish the internal cause of the duplication.

Earlier attempts involving a disabled scenario, and attempts that produced INVALID_EMAIL because of test-email generation, are separately retained as invalid and non-counting. They should not be conflated with the valid boundary-case result.

malformed-website: an invalid value was stored despite normal completion

The string not a URL in the website field was accepted and stored unchanged in HubSpot’s Website URL property. Make completed normally, and the company, contact, and tasks—as well as the welcome email, internal handoff email, and kickoff event—were created.

The observed outcome is propagated. This case shows that the absence of an execution error alone cannot establish input validity. Normal completion is not evidence that the value was rejected or isolated.

repeat-submission: resubmitting identical input increased outputs

This test submitted identical input twice. Both Make executions completed normally. The structured manual-observation record reports two companies, one shared contact, eight tasks in total, two welcome emails, two internal handoff emails, and two kickoff events. The observed outcome is duplicated: in this case, the workflow did not satisfy the property that resubmitting the same input creates no additional unwanted outputs—that is, idempotency.

Only the structured manual-observation record supports these downstream duplicate counts. The saved evidence pack contains no vendor-interface screenshots of duplicated downstream outputs. The saved Make images show only that the two executions succeeded. They are not treated as proof of duplicate counts.

Read the success count separately from the scope that can be delegated

The 19/19 success result is meaningful as confirmation of the defined happy-path processing. At the same time, this set of boundary cases records partial data creation, storage of an invalid value, and duplicated outputs. The happy-path success count alone does not show these behaviors.

When deciding what operational scope can be delegated, it is necessary to assess not only the definition of success, but also where missing inputs are stopped, how the same client ID and resubmissions are handled, and how residual outputs are recovered. These are checks prompted by the observations here; they do not demonstrate behavior after a fix.

The site’s methodology likewise lists recovery from failures and ongoing maintenance alongside reliability in repeated testing. The completion review permits an article that clearly states the limits of its evidence, while setup time remains unmeasured and unrecorded. Having publishable observation records is not the same as being able to draw broad conclusions about reliability or return on investment.

The practical takeaway from this experiment is the behavior confirmed for valid input and the weaknesses confirmed in boundary cases. Reading 19/19 requires considering the denominator, success criteria, separate tests, and evidence limits together.

If you want to build or test a similar automation, you can create a Make account. Affiliate disclosure: this Make link is an affiliate link, and Agency Workflow Bench may earn a commission at no additional cost to you.

Frozen records supporting this article

The following repository references support the counts and behavior in this article. They are saved materials and do not include additional experimental results.

  • Main-test count and treatment of Run 1: benchmark/exp001/results/implementation-a-results.csv, benchmark/exp001/results/implementation-a-main-audit-evidence.json, benchmark/exp001/results/implementation-a-summary.md
  • Frozen definitions of success criteria and duplicate prevention by design: benchmark/exp001/implementation-a-freeze.md
  • missing-email observation: benchmark/exp001/results/edge/observations/edge-missing-email-20260921/missing-email.json
  • Valid duplicate-client observation: benchmark/exp001/results/edge/observations/edge-duplicate-client-20260921-r3-fix/duplicate-client.json
  • Invalid and non-counting attempts: benchmark/exp001/results/edge/invalid-attempts.json
  • malformed-website observation: benchmark/exp001/results/edge/observations/edge-malformed-website-20260921/malformed-website.json
  • repeat-submission observation: benchmark/exp001/results/edge/observations/edge-repeat-submission-20260921/repeat-submission.json
  • Scope of image evidence for the repeated-submission case: benchmark/exp001/evidence/edge/edge-repeat-submission-20260921/README.txt, benchmark/exp001/evidence/manifest.json
  • Conditions for publishing the article and unmeasured items: benchmark/exp001/implementation-a-completion-review.md