Resilience

The Tuesday Problem: What Your Automation Does When the Portal Changes Overnight

Every screen automation is one redesign away from stopping. The question isn't whether the system underneath it will change. It's what your automation does in the minute after it does.

AutomatR TeamOct 7, 202612 min read

On Tuesday 15 November 2022, the Directorate General of Systems at India's Central Board of Indirect Taxes and Customs issued an advisory. ICEGATE — the national portal through which importers, exporters and cargo carriers file their customs declarations — would begin moving to a new website from the next morning. The notice was precise. ICEGATE 2.0 would launch "in phased manner w.e.f. 16.11.2022," location by location, with the personalised dashboard, notifications and the new online filing forms reaching registered users "between 16 November 2022 and 29 December 2022." The old site would stay reachable at old.icegate.gov.in for anyone whose port hadn't switched yet. A ticker on the old homepage said the same thing in one line: from 8 a.m. on the 16th, ICEGATE was migrating.

AdvisoryDirectorate General of Systems · CBIC15 Nov 2022
Subject: Advisory for launch of ICEGATE 2.0

"Directorate General of Systems, CBIC is launching ICEGATE 2.0 in phased manner w.e.f. 16.11.2022 … Users will continue to use existing features of ICEGATE 1.0 and additional features under ICEGATE 2.0 till the complete rollout of the latter."

"… a complete bilingual website which has been designed to provide contemporary user interface for enhanced user experience." · "During online filing, number of fields will be validated simultaneously to user entering values in the forms …"

Phase-I rollout 16 Nov – 29 Dec 2022 by location · ICEGATE 1.0 retained at old.icegate.gov.in · first wave of inland depots and land stations live 16/11/2022; sea ports from 24/11/2022

Nothing went wrong. The advisory was clear, the transition was staggered, the help desk was staffed, and existing users didn't even need to re-register. By the standards of government IT it was a model migration. And on the morning each location switched over, any screen automation built on the old layout and pointed at the new site would have stopped — not because the portal failed, but because the bot had memorised where the buttons were, and the buttons had moved. The old address was a grace period, not a reprieve: every bot had to follow its port to the new site sooner or later, and the day it did, it met screens it had never seen. A "contemporary user interface" is, from a bot's point of view, a different building.

That is the Tuesday Problem.Every system your automation touches will, on some Tuesday, become a different system. (Strictly, the notice came on the Tuesday and the change on the Wednesday. That is the problem in miniature: one day's warning, addressed to people.) The notice will be polite. Your bots will not read it. What matters is what your automation does in the minute after — and there are only three possible answers.

30–50%
Share of initial RPA projects EY said it had seen fail
3% → 4%
Organisations running 50 or more robots, a year apart — the “scale” rate Deloitte surveyed in 2017 and 2018
45%
Firms dealing with bot breakage weekly or more often, in a Forrester Consulting study commissioned by Tricentis
SOURCES: EY, GET READY FOR ROBOTS (2016, PRACTITIONER OBSERVATION) · DELOITTE GLOBAL RPA SURVEY 2017 (400+ RESPONDENTS) AND 2018 · FORRESTER CONSULTING FOR TRICENTIS (FEB 2020, VENDOR-COMMISSIONED)

Read the dates. The newest of those figures is from 2020, the oldest from 2016, and that's the point: the problem has been documented for a decade, the market has bought a great deal more automation since, and nobody has published evidence that it moved. The third figure is from a vendor-commissioned study and says so — a signal, not a measurement. None of them says why bots break, because no primary source we could find measures it. This piece explains the mechanism instead, and ends with a one-minute test you can run yourself.

First, understand what a screen bot actually knows(why RPA bots break when the UI changes)

A screen automation doesn't know what a button is. It knows a description of where one was. That description is called a selector: a path through the page's structure, an element's id, a class name, "the third cell in the second row." When the bot runs, it looks for something on the page that matches the description and acts on it.

When the page changes, one of three things happens to that description. It still matches, because the thing that changed didn't matter. It matches nothing, and the bot stops. Or it matches something else, and the bot carries on — on the wrong element.

The second case is the famous one. "Brittle" is the industry's word for automations that stop whenever a page is redesigned, and it's why bot estates come with maintenance teams. The third case is the one nobody talks about, and the dangerous one: a bot that doesn't know it's lost. A field moves down one row. The selector that said "second row, third cell" now points at a different supplier's invoice. The bot fills it in. Nothing errors. Nobody is told.

So let's name them: the Three Answers

By the end of this piece you'll be able to ask any vendor for a one-minute demonstration that places their automation in one of three categories — and know why the one that never stops is the one to worry about.

Those were the three things that can happen to a selector: it matches, it matches nothing, or it matches something else. The Three Answers are one level up. They describe what the automation is built to do once its selector has failed.

When the screen underneath an automation changes, there are exactly three things the automation can do. Every product on the market, including ours, gives one of these three answers. Most give a different one depending on the step.

ANSWER 01
Break loudly
The selector matches nothing. The run stops with an error. Someone finds out — from the exception count, or from a customer. Classic RPA. Honest, expensive, and the origin of every bot maintenance budget.
ANSWER 02
Heal silently
The system finds “the most similar” element on the new page and carries on. Fast, invisible, and the version most “self-healing” marketing describes. On a payment screen, the most dangerous of the three.
ANSWER 03
Heal, and write it down
The system finds the element by what it is, not where it was; records what it matched, what changed and how sure it was; refuses to guess below a confidence floor; and always stops on a step that can't be undone.
The test

When the screen changed, did your automation stop, carry on, or carry on and tell you what it changed?Only the third answer is one you can audit. The first costs you a maintenance team. The second costs you something you won't find until month-end.

The rollout, one answer at a time

Walk the ICEGATE migration through the three answers. The advisory is on the record; the bots are a thought experiment, not a report of any firm's experience.

01 · Break loudly

The morning a location goes live on 2.0, every declaration bot for that port stops at the first field it can't find. A clearing agent running forty of them finds out from a wall of exceptions, or from a client asking why nothing has been filed. The fix is a person re-recording each selector against the new layout, forty times. This is the maintenance line that never made it into the business case.

02 · Heal silently

The bot finds a field with a similar label on the new dashboard and files. A customs declaration is a legal filing; "similar" is not a standard anyone would accept from a person, and the bot cannot tell anyone it applied one. If it matched wrong, the error is in a government system under the agent's own credentials, and the first sign is a query from customs weeks later.

03 · Heal, and write it down

The bot resolves each field by its properties — label text, role, neighbours — and logs the result: old target, new target, confidence, screenshot. At the submit step, which is irreversible, it stops and shows the card, because an automation that has just adapted to a redesigned page is exactly the one that should not submit a legal filing unsupervised. Someone approves at nine. The log is the change report nobody had to write.

The portal did the same thing to all three. The difference between the second and the third is not whether the automation survived the redesign — both did. It's whether anyone can find out what it did to survive.

Silent healing is older than you think(and the research says it picks the wrong element)

Finding elements by their properties instead of a fixed path is not new, and we'd rather say so than pretend otherwise. The academic work dates to at least 2015 (Leotta and colleagues' multi-locator approaches, then ROBULA+ in 2016). Commercial test tools have described "auto-healing" since at least February 2018, when mabl launched with a feature that, in its words, "automatically heals test [sic] that fail due to small UI changes." Open-source Healenium was public by 2020. In RPA, one of the largest platforms shipped a healing agent in 2025 — and ships it behind a governance switch, with documentation stating that if the setting is disabled, self-healing will not be attempted at execution time. The industry has quietly agreed that healing is something you have to be allowed to do.

Here is what the research says about how well it works. In a 2023 study in ACM Transactions on Software Engineering and Methodology, Nass and colleagues tested a similarity-based locator called Similo across 40 popular websites. It failed to find the correct element in 72 of 598 cases — 12% — against 24% for the baseline. A real improvement, and still one failure in eight. The paper names the risk precisely: above its threshold, the approach "will return a matching but incorrect web element even if the target web element is not yet present." A 2026 follow-up in Empirical Software Engineeringput it in four words — "Similo always returns an element" — and its best hybrid located 98.8% of elements with broken locators. Excellent; and for a locator that always returns something, it still means roughly one heal in eighty lands on the wrong element. Without a log, nobody sees which.

That is the whole argument against Answer 02 in two numbers. Healing works most of the time. The times it doesn't are invisible by design.

Four kinds of change, four kinds of automation(what survives, and what tells you)

Systems change in four ways, roughly in order of how often they happen. Automations answer in four ways. The table reads differently across than it does down, and the twist is in the column that never stops.

Four kinds of change to a target system — a field moves, a label or id changes, a step is added to the flow, the portal is re-platformed — against four kinds of automation: hard-coded selector RPA, an AI assistant driving the screen, property matching with silent heal, and property matching with logged heal and a deliberate stop.
The changeHard-coded selectors(classic RPA)AI assistant on the screen(computer-use agent)Property matching, silent heal(most “self-healing”)Property matching, logged heal(ours — a claim, earned below)
01A field movesStopsOr worse: the positional selector now points at the wrong row.Carries onReads the screen like a person; can't tell you it adapted.Heals, unloggedFinds the field. Nobody knows it moved.Heals, loggedOld target → new target, confidence, screenshot.
02A label or id changesStopsThe id was the selector.Carries onUsually right. Unaudited.Heals, unloggedWhere the wrong-element cases live — one in eight for Similo, one in eighty for its best hybrid.Heals, loggedBelow the confidence floor: stops with the candidates in the reason.
03A step is added to the flowStopsOr skips it, if the next selector still matches.Carries onImprovises the new step. May or may not be the step you'd have chosen.UndefinedNothing to heal to; behaviour depends on the product.Stops by designA step the workflow never contained is not a heal. It's a change request.
04The portal is re-platformedStopsICEGATE, 16 November 2022.Carries onThe strongest case for the assistant — and the one with no record of how.Heals, unloggedSurvives, mostly. Which fields it re-matched wrong, you'll learn later.Stops by designMass re-resolution on an irreversible flow is reviewed, not run.
HEALS, UNLOGGED — THE RUN CONTINUES AND THE CHANGE IS NOT RECORDED  ·  STOPS BY DESIGN — HALTING IS THE INTENDED BEHAVIOUR, NOT A FAILURE

Read across the first row and the classic bot is the only one that stops. Read down the AI-assistant column and every cell is a survival — the column that never stops is the one that never tells you it was lost. Now read the logged-heal column, and notice that two of its cells say Stops by design — the same visible outcome as the classic bot. Be clear about what that costs: on a re-platform, a person still has to rebuild the affected steps, same as with the bot estate you have. What changes is what they start from — a report of which steps moved and where the resolver gave up, instead of forty silent exceptions — and what happened in the meantime, which is nothing irreversible. That is not a weaker product. A pipeline that cannot stop is not resilient; it is unsupervised. The logged-heal column is ours, so treat it as a claim to be tested, not a fact to be read. The next two sections exist to earn it.

Four objections

"We'll just maintain the bots."

You will — and the surveys suggest it costs more than the plan allowed for. Of Deloitte's 2017 respondents who had actually implemented RPA (a small group — 32), 63% said their time-to-implement expectations weren't met; Pega's 2019 survey of 509 decision-makers found 41% saying ongoing bot management took more time and resources than expected. Maintenance is possible. So cost it, with your own numbers in place of ours. Say a back office automates ten portals; each makes a small change — a field moves, a label changes — four times a year; each change costs a person two hours to notice, find and re-record across the workflows it touches. That is 10 × 4 × 2 = 80 hours a year of maintenance that appears nowhere in the business case, before a single re-platform. A logged-heal system absorbs most of the first two rows of the table on its own; what it hands the person is the log, not the hunt. The assumptions are stated so you can replace them. The shape of the sum doesn't change. And the question underneath it is the same: whether the automation tells you it needs maintenance, or waits for a supplier to call.

"Our RPA vendor has self-healing now."

Most do. So ask which answer it gives. Every vendor page says "self-healing"; almost none says what happens when the heal is wrong, whether a heal is logged, or whether healing is permitted on a step that moves money. The Change Drill at the end of this piece takes a minute and answers all three.

"Computer-use agents just look at the screen like a person. Doesn't that solve it?"

Partly, and the benchmarks have moved fast enough that we'll date-stamp this. When OSWorld was published in April 2024, humans completed 72.36% of its desktop tasks and the best model 12.24%. By February 2026 — to take a vendor's own numbers — Anthropic's system card for Claude Opus 4.6 reported 72.7% on OSWorld-Verified, and by May the company had put its successor at 82.3%, a figure it restated after changing how it runs the evaluation "to more accurately reflect the model's performance in the real world." On the task itself, the assistants have caught up. But those are single-attempt runs on a few hundred short tasks; they say nothing about cost per step at ten thousand declarations a month, and nothing about the trail. An agent that "looks at the screen like a person" adapts to a redesign the way a person does — confidently, without writing down what it did differently. A person fills in the wrong field confidently too. What makes either safe is the record.

"A vendor would say this."

Yes. So run the Change Drill on our demo before anyone else's, and watch which of the three answers you get on an irreversible step. If it carries on, we've failed our own test in front of you.

How AutomatR answers the Tuesday Problem

To be plain about status: what follows is the specification our resolver is built to, not a field report. We describe it in terms of what the run log must show, because that is the part you can check in a demo. The principle is the one that runs through this series: the AI designs the workflow, and a deterministic engine runs it. A heal is a deterministic resolver making a logged decision, not a model improvising.

01 · Find
By what it is, not where it was
Each step is captured with several ways of identifying its target — label, role, attributes, neighbours, position — and resolved at run time by scoring candidates against all of them, not by one brittle path. The recorded selector is tried first; the scorer runs only when it fails.
02 · Record
Every heal is written down
When a step resolves to a different element than last time, the log records the old target, the new one, the confidence score and a screenshot. Healing doesn't break the audit trail. It is an audit trail — the most common entry on it.
03 · Refuse
Not sure means stop
Below a confidence floor the resolver doesn't guess; the run stops with the top candidates in the reason. Steps marked irreversible — a submit, a payment, a deletion — never heal silently. They stop and show the card, which is Control 02 from The Off Switch.
Run log · entryStep 07 · supplier codeHeal accepted
old target
#frm > div:eq(3) > input[name="supp_cd"]
new target
label:contains("Supplier code") + input
matched on
label text · role · neighbour "GSTIN" · position (shifted one row)
confidence
0.91 · floor 0.80
before / after
two screenshots attached
written back
yes — next run tries the healed selector first · approver: queue owner role
next step
Step 08 · SUBMIT — flagged irreversible. Run suspended. Card raised.
WHAT A HEAL ENTRY RECORDS. AN ILLUSTRATION OF THE FORMAT, NOT A REAL RUN — THE FIELDS ARE THE DESIGN; THE VALUES ARE INVENTED.
For the engineers on the buying committee — everyone else may skip this section
  1. Capture.At design time each step stores a multi-candidate descriptor: the selector string, the iframe chain, and the element's attributes, visible text, role and geometry.
  2. Resolve. At run time the recorded selector is tried first. On a miss, an in-page scorer ranks candidate elements against the stored descriptor; the top candidate is accepted only above a threshold, and only on steps not flagged irreversible.
  3. Write back. An accepted heal is written back to the step under governance, so the next run tries the healed selector first, and logged with before/after, confidence and screenshot.
  4. Refuse. A rejected heal stops the run; the reason carries the top candidates and their scores, so the person fixing it sees what the resolver saw.
  5. Prior art. The scorer is in the same family as the similarity-based locators cited above. The differentiation is behaviour on failure: a floor, a log, and a rule that irreversible steps never heal unsupervised.

The AI designs the workflow; a deterministic engine runs it.

Three numbers settle a self-healing claim, and they come from run logs, not a brochure: how many screen changes the workflows met in the last twelve months, how many steps were re-resolved and logged, and how many times the resolver chose to stop. Ask us for them in the demo — and ask every other vendor for the same three.

The Change Drill

Run this in any vendor demo, including ours. It takes a minute.

  1. 01Move it.Ask them to move one field on the demo screen — or rename its label — and run the step again.
  2. 02Watch.Did it stop, carry on, or carry on and tell you? That’s the answer, and it took ten seconds.
  3. 03Show me the line.Ask to see the log entry for what just happened. If there isn’t one, you’ve found the silent heal.
  4. 04Now the submit step.Ask them to do the same on the step that can’t be undone, and watch whether it stops. If it does, you’ve found a vendor who knows what “irreversible” means.

A vendor whose automation carries on through step four is selling you Answer 02 under a better name. One whose automation stops is selling you a held filing and a review card: the portal still picks the day, but nothing irreversible happens until a person has looked. That is the Tuesday you can plan for.

Bring us a screen

One process, one screen you expect to change. We'll move the field in front of you.

One last question

The last time a system your automation depends on changed its screen, how did you find out — and how long after?

Frequently asked questions

Why do RPA bots break when a screen changes?

A screen automation doesn't know what a button is; it knows a description of where one was — a selector. When the page changes, the selector either still matches, matches nothing (the bot stops), or matches something else (the bot carries on, on the wrong element). The third case is the dangerous one because nothing errors.

What should self-healing automation actually mean?

Finding elements by their properties rather than a fixed path, recording every heal with before, after and confidence, refusing to guess below a confidence floor, and never healing silently on an irreversible step such as a submit or a payment. Healing that isn't logged is an unrecorded change to a production process.

Don't computer-use AI agents solve this by looking at the screen like a person?

Partly. As of October 2026, vendor-reported scores on OSWorld-Verified exceed the 72.36% human baseline. But those are single-attempt runs on a few hundred short tasks, and an agent that adapts to a redesign like a person does so without a record of what it did differently. What makes either safe is the trail.

How do I test a vendor's self-healing claim in a demo?

Run the Change Drill: ask them to move or rename one field on the demo screen and run the step again; watch whether it stops, carries on, or carries on and tells you; ask to see the log line; then repeat on the submit step and watch whether it stops.

Further reading: going deeper on resilience

SourcesICEGATE 2.0: "Advisory for launch of ICEGATE 2.0," Directorate General of Systems, CBIC, 15 November 2022 (a Tuesday) — phased launch "w.e.f. 16.11.2022"; dashboard, notifications and web forms rolled out to registered users "between 16 November 2022 and 29 December 2022" by location; ICEGATE 1.0 retained at old.icegate.gov.in; quoted wording as in the advisory. "ICEGATE 2.0 live locations," CBIC, November 2022: inland container depots and land customs stations live with 16/11/2022, sea ports from 24/11/2022. Homepage ticker, old.icegate.gov.in: "From 8 AM IST, 16-Nov-2022 ICEGATE will be migrating to new website." Existing users not required to re-register: ICEGATE SEZ Registration advisory, 4 August 2023. The old portal was not formally shut down on a date we could find; services have been withdrawn piecemeal since. What bots did on the day is reasoning about the documented change, not a report of any firm's experience. EY, "Get ready for robots: Why planning makes the difference between success and disappointment," 2016, p. 2: "we have seen as many as 30 to 50% of initial RPA projects fail" — a practitioner observation, not a survey. Deloitte, "The robots are ready. Are you?" (Global RPA Survey, September 2017, over 400 respondents): "only 3% of organizations have managed to scale RPA to a level of 50 or more robots"; "63% said their expectations of time to implement were not met" is drawn from the 32 respondents who had implemented. Deloitte, "The robots are waiting," 2018: "Only four per cent of respondents to our survey are operating more than 50 robots." Forrester Consulting study commissioned by Tricentis, "Barriers and Best Practices for Scaling RPA," announced 12 February 2020: "Forty-five percent of firms deal with bot breakage on a weekly basis or more often"; sample size not published. Pegasystems survey, 10 September 2019, 509 decision-makers: 87% experience some level of bot failures; 41% say ongoing bot management takes more time and resources than expected. Both vendor-commissioned. No primary source was found for any figure on the share of RPA cost that is maintenance or the share of failures caused by UI change; none is quoted. Prior art: Leotta et al., multi-locator web testing (ICST 2015) and ROBULA+ (2016); mabl launch, 21 February 2018 ("mabl automatically heals test that fail due to small UI changes"); Healenium public by 2020; a major RPA platform's healing agent, generally available 2025, and its governance setting, paraphrased from that vendor's public documentation (vendor not named by editorial policy). Nass, Alégroth, Feldt, Leotta & Ricca, ACM TOSEM 32(3), 2023: Similo failed in 72 of 598 cases (12%) against 146 (24%) for the baseline; "will return a matching but incorrect web element even if the target web element is not yet present." Kluge & Stocco, Empirical Software Engineering 31:175, 2026: "Similo always returns an element"; VON Similo "tends to produce more false positives than Similo"; HybridSimilo "locates 98.8% of elements with broken locators." OSWorld: Xie et al., arXiv:2404.07972 (April 2024), humans 72.36%, best model 12.24%; Anthropic, Claude Opus 4.6 System Card (February 2026), OSWorld-Verified 72.7%; Anthropic, Claude Opus 4.8 announcement (May 2026), footnote: the Opus 4.7 OSWorld-Verified score updated to 82.3% after "changes to how we run the OSWorld-Verified evaluation in order to more accurately reflect the model's performance in the real world." Benchmarks move monthly; figures are as of October 2026, when this was written. AutomatR's resolver is described as designed: the multi-candidate capture, weighted scoring, confidence floor, governed write-back and irreversible-step rule are the specification the engine is built to; the three run-log figures are for the demo, not quoted here. The heal-log entry shown is an illustration of the record format; its values are invented.
AutomatR
AutomatR TeamAutomatR — Unified Agentic Stack for Enterprises