Crisis Indecision
Here's a question that stopped a boardroom cold last year:
"If your organisation had to operate without IT for 72 hours starting right now — what would actually break first?"
Not the systems. Not the servers. It’s more likely that decision making will be impacted the most.
After many years of helping organisations develop business continuity plans, improve crisis management and working with Boards, I've noticed something consistent. Organisations invest heavily in recovery plans but chronically underinvest in crisis decision-making capacity.
Plans sit in folders. Exercises happen once a year if you're lucky. And when a real event hits - a cyber incident, a supply chain failure, an extreme weather event - the bottleneck is almost never technical. It's the leadership team not knowing who has the delegated authority, who communicates what and in what order. Often, reporting obligations to regulators are unknown.
The frequency and intensity of disruptions is increasing and many Boards are giving insufficient attention to resilience.
The Boards I've seen handle crises well share three things:
· They've practised decision-making under pressure, not just plan-reading
· They have clear escalation protocols the whole organisation understands
· They treat resilience as a strategic asset, not a compliance checkbox
Governance isn't just about what's on the agenda. It's about whether your organisation can keep functioning - and keep serving its stakeholders when things go wrong.
What's your organisation actually exercising? And how confident are you that your Board could navigate a 72-hour disruption?
Minding the gap
Something I've observed consistently across more than 30 years in business continuity and resilience consulting:
Boards are getting better at asking about risk. They're not always getting better at understanding it.
There's a meaningful difference between a board that ticks the risk register at each meeting and one that genuinely understands the interdependencies that make a threat material — or not.
I see this most clearly now in two areas: climate risk and cyber security.
Both have become fixtures on board agendas. Both attract well-meaning attention and considerable reporting. But too often, the conversation stays at the surface — headline risks, regulatory obligations, general frameworks.
What's missing is operational depth. The board asks "are we exposed to cyber risk?" when the more important question is "do we understand our critical business processes well enough to know which systems, if disrupted, would genuinely threaten our viability?"
Resilience isn't a compliance outcome. It's a strategic capability. And boards that treat it as such — by investing in genuine understanding of how their organisation actually functions under stress — make qualitatively better decisions when it matters most.
The organisations that navigate crises well don't discover their gaps during the crisis. They've already mapped them.
How is your board approaching the gap between risk awareness and risk understanding?
Business Continuity inaction
On 8 July 2026, Telstra’s national mobile network collapsed after a faulty software update reset a GPS‑based Network Time Protocol server back to 2006. That single failure corrupted the clock that synchronises signalling, authentication and routing across the entire mobile network. As bad timestamps rippled through the system, network nodes fell out of sync, triggering a nationwide shutdown of voice, data and signalling.
The outage exposed deeper weaknesses: Telstra had relied on legacy mid‑2000s time‑keeping infrastructure, ignored repeated warnings about GPS timing vulnerabilities, and operated a single point of failure with no effective disaster‑recovery backup. Even the regulated “camp‑on” mechanism for Triple Zero failed to fully transfer calls to alternative carriers.
Telstra CEO Vicky Brady appears before a Senate inquiry on 17 July 2026 into the Telstra outage.
V/Line was hit hardest. Its regional trains depend entirely on Telstra’s 4G network for safety‑critical radio communication. When the network went dark, V/Line had no alternative system to keep trains operating and was legally required to halt all services. Recovery was slow and painstaking. Telstra and ARTC first had to stabilise network timing, then V/Line conducted train‑by‑train radio testing to prove communications were continuous and reliable before any service could resume.
As digital‑economy expert Paul Budde noted, Australia must stop reacting to outages and start designing resilient telecommunications systems before failures occur.
Most Australian organisations are smaller than Telstra and not as complex. However, this incident is a salutary reminder that even very well resourced organisations are vulnerable to failures if management doesn’t pay sufficient attention to their own procedures.
See here for the full article from Paul Budde:
Outsourcing IT Operations
We recently attended a very interesting seminar organised by the Global Association of Risk Professionals (GARP). The topic was to help finance companies address APRA’s CPS 230 standard that has come into force on 1 July 2025. The standard addresses the need for finance companies to develop business continuity plans and to prudently manage their operational risks. In light of the common approach to outsource IT operations to third parties, the standard focuses heavily on this trend.
Speakers were from UniSuper, NAB and Deloitte and they spoke generally about their experience in preparing for the introduction of the new standard.
The speaker from UniSuper focused on the company’s experience of having Google inadvertently delete the entire UniSuper Google Cloud subscription, impacting over 600,000 customers. This occurred even though UniSuper had duplicate infrastructure and data in two geographies.
The conversation settled primarily on the challenge of maintaining the resilience of IT applications and data, in light of the small number of vendors in Australia.
See below for a diagram that depicts the major vendors of IT infrastructure and applications in Australia.
By way of example, before cloud and SaaS, the four banks in Australia would own and operate their own data centres, IT systems, networks and purchase the application software licences to run their businesses. If one of the banks suffered a power outage, a fire in their data centre or an IT malfunction, only its customers were impacted.
Today, it is likely that all four banks subscribe to the services of AWS, Microsoft and Google. So, if one of these providers suffers and outage, many more bank customers could be impacted.
The dominance of Microsoft is particularly concerning because it operates cloud services and three dominant SaaS services – O365, Teams and SharePoint. For a large proportion of Australian organisations, employees working from home are especially reliant on Teams.
The other aggregation of risks results from the concentration of data centres in Melbourne. There are currently four large data centres located in close proximity to each other in Port Melbourne, with NextDC planning another very large data centre nearby. Port Melbourne is about 2-3 metres above the Yarra River, which is open to the sea.
Uncertainty in the US
Finally, Jeff Bezos and Mark Zuckerberg have recently substantially changed their policies governing The Washington Post and Facebook. It’s plausible that given the major changes occuring in the US, that other companies mentioned above in the diagram, could also initiate substantial changes to the way they operate, possibly impacting Australian companies.
What we used to take for granted is no longer!
We left the seminar believing that more Australian organisations should seriously consider APRA’s approach to managing their reliance on IT systems and data.
The introduction of the Standard on 1 July 2025 adds urgency for Australia’s regulated entities!
PS: The incredibly impactful electrical sub-station fire at Heathrow Airport recently apparently also supplied a number of the UK’s data centres!
Floods and Resilience
Reducing the impact of a flood
We often advise clients that the risks presented by climate change are increasing rapidly. For any organisation that has assets exposed to flooding, sea level rise or storms, mitigating these risks can be challenging. Here are a couple of success stories.
In May 2010, a major flood hit Coca-Cola’s 30,000 m2 bottling plant in Nashville, Tennessee. The facility is located in a high hazard, 100 year flood zone.
The flooded bottling facility during the 2010 flood which prompted the flood mitigation project.
Coca-Cola partnered with FM Global to develop a plan to protect their facility from future floods.
They decided they could not relocate the very large warehouse away from a known flood risk, so they developed a method to reduce the impact of the next inevitable flood. They protected critical production equipment within the facility using flood walls. This enabled them to let the flood waters flow into, and out of, the building.
Flood wall around the critical infrastructure and the flood door to allow the water to flow back out of the building.
Amazingly, they were able to verify the effectiveness of the solution during another serious flood in March 2021. See here for more details: https://www.fm.com/insights/coca-cola
Interestingly, the Reject Shop 26,000 m2 Distribution Centre in Ipswich, Queensland had a very similar experience. After the extensive flooding at the start of 2011, they installed a floodbarrier system around the DC. Again, they had the opportunity to test the barrier when another flood hit the area in 2013. There was no impact to the DC’s operations!
Resilience in the Cloud
The Uptime Institute recently published an excellent paper on the topic of the cost and benefits of purchasing increased resilience from cloud providers, using AWS as a case study. The baseline comparison was a system with no resilience installed. The study shows the cost of the increased resilience and the associated reduced downtime.
The author provides some wise advice:
"Unlike privately owned and co-located data centres, customers using the public cloud have no visibility or control over the datacentre used by their cloud provider. When architecting a cloud application, it is up to software developers to incorporate resiliency into their application architecture. Conversely, in amore traditional non-cloud application, data centre teams, infrastructure engineers and software developers should work together to meet resiliency requirements.”
"If customers use more resources to architect resiliency, they need to pay for those additional resources. The implication is that resiliency is neither included as standard nor guaranteed. Customers should design their applications to meet availability requirements and balance this objective against the cost."
As CIO’s are relinquishing control over the operation of their IT systems, it is imperative that the resilience requirements of each application and its data are fully specified in the service agreement with the cloud provider. APRA’s standard soon to come into effect, CPS 230, addresses many of the issues associated with outsourcing critical services to Third Parties. Source: https://intelligence.uptimeinstitute.com/resource/cloud-availability-comes-price
Operational Risk Management - CPS 230
The Australian Prudential Regulation Authority (APRA) has released a guide that covers the new standard on Operational Risk Management - CPS 230. The standard came into force this month.
Although APRA’s standards are intended for companies operating in the Australian financial market, we think the standard and guide provide very good advice for most organisations that are concerned about their operational resilience.
The standard addresses the following:
The assessment and management of a wide range of operational risks, including legal, regulatory, compliance, conduct, technology, data and change management risks.
Business continuity and how organisations should identify time critical business activities and estimate their tolerance for having them unavailable. Importantly, the business continuity plan should document the recovery procedures and workarounds if any supporting resources (people, facilities and IT systems) become unavailable because of a disruption.
Development of a policy for dealing with material service providers. This policy should cover how to identify, manage and monitor the service providers that have a significant impact on the organisation’s operations. They should also evaluate the risks posed by these service providers, sign formal contracts with them, track their performance and carefully manage any major changes in their arrangements.
Managing outsourced IT services
We find that many organisations have outsourced large parts of their IT infrastructure to service providers and as a result they have often yielded management and control to others.
This makes it challenging for the CIO to ensure that the recoverability of IT systems meet the needs of the business. Some Software as a Service vendors will not warrant a Recovery Time Objective. Often, the outsourced system (and its data) only exists at one location, making it a single point of failure.
It is critical that business management identifies the time critical activities and their tolerance to disruption. These requirements should be communicated to IT management, so that the critical IT systems have the necessary resilience to support the business during disruptions.
CPS 230 outlines an excellent approach to achieve that!
DP World cyber-attack
Photo: DP World
Late Friday 10 November, Australia’s largest port operator, DP World, suffered a serious cyber-attack. To minimise the impact of the attack, it shut down its connection to the internet, causing considerable disruption to port operations. It is gradually restoring operations, but could suffer additional pain due to industrial action.
“Even if DP World recovers from the cyberattack to full operations shortly, GuardianAustralia understands customers remain frustrated at the prospect of delays due to protected industrial action from dock workers in coming days.”
APRA getting serious on cyber
CRN reports that APRA is losing patience with regulated entities:
"Three years ago, APRA’s information security standard CPS 234 came into force, and yet many entities are still struggling with foundational issues: ensuring third party controls are effective, making sure that systematic security control testing is in place, and regularly testing incident response plans," APRA Chair Lonsdale said.
"With the potential for serious impact to millions of Australians, our patience has run out."
Don’t forget that APRA not only focuses on regulated entities such as banks, insurers and superannuation companies, but also on the suppliers of material services to these entities.
In July, APRA also announced the new standard “CPS 230 Operational Risk Management”.
Chair John Lonsdale said “We expect regulated entities to be proactive in preparing for implementation, rather than waiting until the last minute to get ready to meet the new requirements. There will be a transition phase for existing contractual arrangements with material service providers for entities that need some flexibility.”
The key to becoming CPS 230 compliant is to start now! There is a lot to do and July 2025 will come around very quickly.
Lessons learnt from the Optus outage
We reckon that there are two big lessons for all of us that have resulted from the Optus outage.
Lesson 1 - Our dependence on functioning telephony and data networks
The breadth of the impact has been enormous and our dependence on all things digital will increase over time. The use of multi-factor authentication to verify your identity using your mobile phone made the outage even more impactful. Here are useful ideas we have collected in the last few days which will help reduce the impact of the next outage:
Paul Budde’s recommendation that we should be able to roam on the network we normally do not use (pretend you’re overseas during the outage of your BAU provider).
Great Vox article on ensuring you can survive if you lose access to your mobile.
Lesson 2 – Good crisis communications is critical
The importance of clear and effective crisis communications is essential to maintaining your good reputation. Even though Optus last year was subject to a major cyber-attack that impacted many customers, it has not improved its ability to communicate. Madonna King neatly sums up the essentials of good crisis communications management:
Communicate quickly and stay on the front foot.
Be clear and don’t sell.
Be sincere.
Be mindful of the impact on your stakeholders.
Incredibly, telcos are currently not part of the Federal Government’s Security of Critical Infrastructure Act (SoCI) !
Another interesting AFR article comparing the treatment of Optus and DP world in the press.
Cyber Security priorities and investments with an outcome-driven approach | The Reboot Show and TrustedImpact
Often when I ask an executive if a service provider's Continuity Plan has been practised, they don't know, which is worse than having no plan. Ben Scheltus, General Manager, Continuity Matters
Readiness to detect, contain and respond to Information Security threats is measured by an organisation's state of Cyber Maturity. The Cyber Maturity journey requires strong leadership direction and sustained action - it's not a software procurement matter that can be left solely in the hands of IT departments or Information Security generalists.
An organisation's state of Cyber Maturity, at any point in time, determines its level of Cyber Resilience - which is the organisation's ability to recover from a cyber crisis when it happens.
The Reboot Show, in conjunction with TrustedImpact hosted a series of leadership discussions, for executives and board members, with 9 Cyber Security experts in Australia to unpack modern security perspectives and reflect on contemporary misconceptions.
This discussion paper summarises key insights shared by 9 Cyber Security experts including:
Executive responsibility for preventable crises
The Cyber Maturity continuum and building Cyber Resilience
Navigating business risks at the speed of software
Unique risks associated with cloud services
Creating engagement through training and awareness
TrustedImpact's Cyber Security Training and Awareness Program Pillars
Limitations of Penetration Testing
A guide to the TCFD
The casualties of the climate crisis could include financial stability, the global economy, and the value of investments. As governments catch up to the realities of climate change and the policy response continues to gather pace, global markets need transparency into the financial impacts of climate change on companies.
The Task Force on Climate-related Financial Disclosures (TCFD) released their recommendations in 2017 to improve and increase reporting of climate-related financial information. Today, 2,000+ organizations support TCFD, including 110+ regulators and government entities across 78 countries.
Why read this guide?
It outlines the benefits of climate reporting to firms such as yours, and explains how companies and regulators are implementing the TCFD recommendations.
Firms implementing the recommendations are able to:
Efficiently identify climate-related opportunities and risks
Proactively address investors' demands for climate-related information in a framework that investors are increasingly asking for
More effectively meet current requirements to report material information in financial filings
Enhance risk management and strategic planning, through better understanding of climate risk
Bloomberg has created this guide to help you better understand the benefits of implementing the TCFD recommendations.
Global silicon chip shortage hits supply of phones, TVs, cars and Australia's NBN | The Guardian
A global shortage of one crucial piece of technology is causing delays in everything from cars and televisions to video game consoles and Australia’s National Broadband Network rollout.
A global shortage of one crucial piece of technology is causing delays in everything from cars and televisions to video game consoles and Australia’s National Broadband Network rollout.
Giant ship blocking Suez canal partially refloated | The Guardian
One of the largest container ships in the world has been partially refloated after it ran aground in the Suez canal, causing a huge jam of vessels at either end of the vital international trade artery.
The Security Leadership Series | TrustedImpact and AISA
The very rapid adoption of cloud is exposing organisations to increased risks as they generally exercise an overabundance of trust and a scarcity of caution.
Read this very informative paper by TrustedImpact and AISA on how you can better manage your information that is held by your cloud provider.
Texas Cold Crisis: Insurance Options for Severe Weather Disruption | Risk Management Monitor
On February 15, a massive and unseasonal storm with frigid temperatures spiked the demand for power and outpaced the supply, severing power to 26 million Texans. Unpredictable weather patterns present risks for business owners, but also create an opportunity to improve their risk mitigation strategies to address future uncertainties.
Power outages are not caused by storms alone. Heat waves, hurricanes and wildfires can also create power outages—and outages are more common than business leaders may think. S&C’s 2018 Commercial and Industrial Power Reliability Report found that one in four businesses experience at least one power outage per month.