# Mohd Atasha: full site text Strategic advice for chief executives and their teams on the decisions that change a company, working across business model, digital technology, policy and approvals, leadership and change management. Increasingly that means deciding which calls a person should make and which can be left to a machine. Better decisions are what produce growth, resilience and trust. Foreign direct investment negotiated for the Malaysian government, institutional operations run across Asia and Europe, and a contribution to national AI technical standards. Risk work covers air gap to zero trust, on-premise to cloud, and deterministic systems to AI; it is one part of that, not the whole of it. ## Standing ### AI Security and Resilience subgroup, Artificial Intelligence Standards Task Force 2025 to present. Convened by the Malaysian Technical Standards Forum, the body designated by the Malaysian Communications and Multimedia Commission for communications and multimedia standards. The task force develops Malaysia’s national AI technical standards for that industry. A technical code on artificial intelligence cybersecurity architecture requirements is listed by the forum as under development. Source: Task force announcement, MTSFB, https://mtsfb.org.my/mtsfb-kicks-off-ai-standards-task-force-to-drive-national-ai-standardisation/ Basis: What a technical code is, MTSFB, https://mtsfb.org.my/technical-code/ ### Contributor, Artificial Intelligence Systems Cyber Security Framework (AISCF) Taskforce 2025 to present. Published by the National Cyber Security Agency and launched by the Prime Minister at the National Cyber Security Summit in Putrajaya on 9 July 2026, together with the National Cryptography Policy. The AISCF sets out how to secure an AI system across its whole life, from inception and design through deployment, operation and retirement, and it protects three things in layers: AI data, AI models, and AI infrastructure and applications. It is addressed to any organisation that builds, supplies or uses AI systems, in the public sector or the private one, and it is aligned with the Cyber Security Act 2024. What it offers is a structure for deciding which controls a given AI system warrants. Source: Download the AISCF from NACSA. Contributors are acknowledged on page 58, https://nacsa.gov.my/aiscf.php Basis: Cyber Security Act 2024 (Act 854), Attorney General’s Chambers, https://lom.agc.gov.my/act-detail.php?language=BI&act=854 ### Partner Advisor, eFounders Fellowship 2018 to 2023. Appointed by Alibaba Group and the United Nations Conference on Trade and Development. Curriculum design for a programme delivered at Alibaba Business School, whose 2018 letter of invitation records him as an advisor of the course. Source: Advisor listing published by Alibaba Group, https://activity.alibaba.com/supplier/eFounders-Partners-Atasha-Alias.html ### Participant, UNCTAD Multi-year Expert Meeting on Investment, Innovation and Entrepreneurship 2020. Eighth session, held at the Palais des Nations, Geneva, on 21 September 2020. He appears on the official list of participants, United Nations document TD/B/C.II/MEM.4/INF.8. Source: List of participants (TD/B/C.II/MEM.4/INF.8), UNCTAD, https://unctad.org/system/files/official-document/ciimem4inf8_en.pdf ### International Consultant, International Trade Centre, Geneva 2019. Netherlands Trust Fund IV, the export sector competitiveness programme run by the International Trade Centre with the Dutch Centre for the Promotion of Imports, which ran from September 2017 to June 2021. The International Trade Centre is the joint agency of the World Trade Organization and the United Nations. The source below establishes the programme, its funder and its dates; it does not name him, because the centre does not publish the names of consultants engaged on a programme of this kind. Source: Netherlands Trust Fund IV programme, International Trade Centre, https://www.intracen.org/our-work/projects/netherlands-trust-fund-iv-export-sector-competitiveness-programme-in-it-ites ## Doctrine ### From air gap to zero trust Version 1.1, revised 2026-07-27. An air gap removed the path by which anything inside could be discovered, reached or administered. Connectivity did not have to cost you that last part; opening inbound ports did. Plain summary: Some networks were built to be completely cut off from the internet. That separation is called an air gap. It worked because there was no route from outside to anything inside. Over time those networks were connected anyway. A supplier needed to log in for maintenance. A new sensor needed to send readings out. Each change was small and each was approved. Together they mean the separation no longer holds. An air gap gave a plant several kinds of protection at once, and most of them cannot be recovered once the cable is in. One can. Nothing outside needs to be able to open a connection inwards. This document argues the mistake was the direction of travel. A door that opens inward can be found and pushed. A connection that only reaches outward gives an outsider nothing to find. You can have the access you need without leaving the plant itself exposed. #### What an air gap actually delivered The air gap is usually described as a physical fact: two networks, no cable between them. The description is accurate, and earlier revisions of this document went too far in dismissing it. Separation is precisely what an air gap was, and it is worth being honest about how much it bought, because the argument that follows only needs one part of it and should not pretend to restore the rest. What it removed was the normal network path by which a remote service could be discovered, reached and administered. Several distinct protections travelled with that. There was no routable path across the boundary. There were no reachable trust relationships to abuse, no management plane an outsider could log into, no ordinary remote administration to hijack, and no internet-borne delivery route for a payload. In many real plants it also meant no externally reachable listener existed inside the OT environment at all, though a machine in an air-gapped room may well have had services listening on its local interfaces. The difference was that nothing outside could reach them. Three things are worth keeping apart, because conflating them is what makes this conversation go wrong. Air-gapped isolation means there is no normal network path across the boundary. No directly reachable inbound service means no external route or listener is exposed into the protected environment. Zero trust means access is authenticated and authorised explicitly, per resource, with no implicit trust granted by network position. Those can coexist and they are not interchangeable, and this document is about the second one. That is the distinction the whole argument rests on. Isolation was one way to arrive at the second property. It was never the only way, and it was always the most expensive way, because it also removed the diagnostics, the telemetry and the vendor support that the plant genuinely needed. Every organisation that later connected its OT network was, in effect, deciding that the price of the mechanism had become higher than the value of everything it delivered. That decision was often correct. What went wrong is narrower and more recoverable than the loss of the air gap in its entirety: when connectivity arrived, reachability was reintroduced through inbound remote-access services, and nobody recorded surrendering that as a decision at all. The joint guidance of 29 April 2026 on adapting zero trust to operational technology holds both halves of this in the same document, which is why it repays a close reading rather than a summary. It notes that legacy OT already possesses foundational segmentation techniques, "including physical air gaps, dedicated virtual local area networks (VLANs), and carefully controlled access to sensitive zones, providing a useful baseline for ZT implementation", ZT being its abbreviation for zero trust. Twelve pages earlier it warns operators to "avoid procuring components or systems that assume security through air-gapping or segmented architecture alone, as modern threats can exploit the false sense of isolation these models provide". Those two statements only look contradictory if you treat the air gap as a mechanism. Read as a property, they are consistent: the property was real and worth keeping, and assuming you still have it because you once did is the error. #### How it eroded, one approved change at a time No plant lost its air gap in a single change. It went in increments, each one documented, justified and signed. A turbine vendor needed maintenance access under a support contract the plant could not operate without. A regulator or a corporate reporting line needed process data upstream, so a historian was given a route out and then, for reconciliation, a route back in. An integrator arrived for a commissioning window with a laptop and a deadline. A remote pumping station forty kilometres away had to be reachable at three in the morning by somebody who was not going to drive there. The joint guidance states the resulting bind in one sentence, and it is worth quoting because it concedes both sides: "Remote access is a major weakness in OT as it represents an initial access vector into an insecure legacy network, but remote access may be necessary for operating distributed infrastructure." Necessary and dangerous, in the same breath. Nothing in that sentence can be resolved by telling an operator to be more disciplined. The mechanism of the erosion is administrative, not technical. Each change was assessed against the question "is this access justified", which it invariably was, and never against the question "what does this leave listening when the work is finished". The first question belongs to change control and has an owner. The second belongs to architecture and, in most institutions, has none. So the plant accumulated inbound pathways the way a building accumulates keys, and the total was never anybody’s document. Ask for the list and you will usually be given a firewall ruleset, which is a record of what was permitted rather than a record of what is reachable. #### Inbound is the thing that cost you, not connectivity Connectivity is a word without a direction in it, and that is why it takes the blame. Almost every access requirement in the previous section can be met by a connection the protected machine opens itself. Almost every one of them was instead met by a connection something outside opens to the protected machine: a forwarded port, a published address, a VPN concentrator, a listening service. The requirement was access. The implementation was exposure. They are not the same thing and the difference is directional. The reason direction matters more than protocol is that it changes what an adversary has to do. Against an inbound service, reconnaissance is free and continuous: the port is there whether or not anyone is attacking it, and it can be found by somebody who has never heard of the plant. Against an outbound-only endpoint there is nothing to enumerate, so the adversary needs a foothold somewhere in the path before the target is even visible. That does not make the system safe. It moves the work from scanning, which is cheap and automatable, to intrusion, which is neither. This is the sharpest available version of the claim, and it is not this practice’s invention. The joint guidance sets out the difference between IT and OT segmentation in a table, and under the row headed Directionality the OT column reads: "Primarily unidirectional–OT systems push data out, minimizing inbound connections". A federal joint publication, aligned to the NIST Cybersecurity Framework functions and to ISA/IEC 62443, states the principle plainly. What the guidance does not do, and what the rest of this document attempts, is follow it to its conclusion for the access cases that appear to require an inbound path. #### Air-gap zero trust network access: keeping the property after the wall goes Air-gap zero trust network access is the name this practice gives to the architecture that follows from the previous section, and it is a local analytical term rather than a recognised category: nobody should go looking for a product sold under it. The protected machine holds no listening service and no published address. A connector inside the protected environment establishes and maintains an authenticated outbound channel, and authorised sessions are then brokered or multiplexed through it according to the design of whatever product is doing the work. The connector-to-broker channel wants strong mutual workload authentication, commonly mutual TLS with rotated certificates or hardware-backed keys, while the human at the other end is separately authenticated through multifactor and privileged-access controls with time-bounded policy. Those are the right patterns rather than inherent properties of the arrangement, and describing them as inherent was an overstatement in earlier revisions. The honest version of the payoff, stated before the diagram rather than after it. An adversary scanning the plant's externally reachable address space finds no directly exposed OT service, which is the part of the air gap's protection that can survive connectivity. What they do find is the broker or service edge, because the externally reachable surface has moved there rather than ceased to exist. That component must be designed, monitored and operated as a high-value control-plane component, and the trade only makes sense if it is. What the plant gets in exchange is that the thing anybody can reach is now a component it chose, placed, monitors and can replace, instead of a listener on the boundary of a plant that cannot be patched during production. The vendor still gets their maintenance window. The distinction from mainstream zero trust network access is the part that matters, and it is where a sceptical reader should push. This section used to draw it wrongly, and the correction belongs in the open, because getting it wrong would cost the argument its credibility with exactly the readers it needs. The claim was that conventional ZTNA puts a gateway in front, so the client reaches the gateway and something on the protected side still listens. That describes a 2020-era deployment and it is not what mainstream ZTNA does now. The NCSC's own reference architectures set out the modern pattern plainly: "At the edge of each network segment, a component designed to be exposed to the public internet such as a connector, proxy or a VPN endpoint is deployed and will establish outbound (reverse) tunnels to the policy engine", which allows "access to be brokered without requiring inbound network connectivity to the application or hosting environment". Outbound initiation is not this practice's discovery. It is the reference architecture. The honest distinction is therefore narrower than the one this document used to draw, and it is not about the direction of the first move. Both patterns remove the directly reachable listener from the protected environment, and both therefore relocate the reachable surface rather than remove it, onto a connector, policy engine or provider edge that is deliberately exposed and has to be run as a high-value control-plane component. What stays genuinely different for operational technology is where policy is enforced, which component remains publicly exposed and who operates it, how sessions are brokered and cut, and whether the arrangement tolerates the timing, safety and vendor-support constraints of a plant that cannot be rebooted to replace a certificate. Those are the questions to press. The history of the last several years is unkind on the exposed component in either design: the concentrators, gateways and edge appliances bought to reduce exposure have themselves been a recurring initial access vector, for the ordinary reason that a device which must accept connections from everywhere is reachable from everywhere. That argument survives the correction; it simply applies to the broker here too, which is the subject of the last section. There is hardware precedent for the direction argument and it is worth being clear about its limits. The joint guidance lists data diodes among OT segmentation tools and describes them as enforcing "unidirectional communications via hardware controls". A diode is the strongest possible statement of the principle: one direction, no return path, physically. That is why it suits telemetry and historian replication and why it cannot serve a maintenance session. The argument here is narrower and more useful in brownfield: take the interactive session an operator actually needs, and carry it over a connection the protected side initiated, so that the direction of the initiation is decoupled from the direction of the work. The honest tension with the guidance sits here rather than being tucked into a footnote. Having stated that OT should be primarily unidirectional, the same document later recommends a jump host: "Jump hosts are dedicated, hardened jump boxes within the OT demilitarized zone (DMZ) acting as the sole entry point for remote access. The authoring agencies strongly recommend a jump host for adding user authentication to legacy networks and enforcing segmentation." A jump host is a listening service. Concentrating remote access into one hardened, monitored, multifactor-protected entry point is a very large improvement on the alternative, which is per-device exposure scattered across a plant, and for an operator who cannot change their architecture this year it is the right advice. What it is not is a restoration of the property, and the guidance’s own Directionality row is the reason why. Both things are true, and a document that pretended otherwise would be less useful to the person who has to choose. #### What the joint guidance of 29 April 2026 actually asks for The guidance is organised on the Cybersecurity Framework functions and is careful to position zero trust as working alongside safety rather than above it. Read as an operator with equipment older than the document, three of its asks are immediately actionable, and several quietly assume a plant you do not have. What is realistic. The asset inventory is genuinely first, genuinely foundational, and achievable in a brownfield plant if you accept passive discovery. The guidance is unusually candid about why that caveat exists, noting that active scanning "may knock a legacy device offline", so it asks for comprehensive inventory while conceding that the standard IT means of obtaining one is unsafe in the environment where it is being asked for. That is not a contradiction, it is a tooling constraint, and it is the single most useful sentence in the document for anyone building a business case: budget for OT-aware passive discovery, not for a scanner licence. Segmentation as policy rather than architecture is also realistic and underrated. The instruction to treat segmentation as "a dynamic, enforceable security policy instead of a one-time architectural decision", and to manage it out of band of the operational network, is achievable without touching a single controller. Where it assumes a greenfield. The identity sections are the clearest case, and the guidance says so itself: "Many OT systems predate modern ICAM capabilities and often require compensating controls above the device level." Read that carefully. The identity layer cannot reach the thing being protected, so authentication happens somewhere upstream of the asset and the asset continues to trust whatever arrives. Every claim about per-asset identity in an OT environment should be read against that admission. The same pattern recurs in secure communications: the modern protocol variants exist, and the document notes they are "often disabled to allow for simpler integration or backwards compatibility", with the fallback being to wrap legacy protocols in TLS-enabled gateways so that "encryption and authentication occur outside the control devices themselves". That is the brownfield answer, and it is worth naming plainly: you are not securing the protocol, you are building a shell around a protocol that cannot be secured. Two further asks deserve credit for realism rather than criticism. Emergency access is treated as "non-negotiable", with break-glass accounts held to limited lifespans and stringent auditing, which is a guidance document conceding that a control which can prevent an operator from reaching a plant in an emergency will be removed by the operator, correctly. And on patching it declines the usual answer: patch windows are operational, patching outside them "is discouraged unless there is an outsized risk of exploitation and impact", decisions are pushed towards a structured method such as SSVC, and where patching cannot happen promptly, "The focus should shift to isolating vulnerable systems and applying compensating controls". An OT operator has heard "just patch it" for twenty years. This document does not say it, which is the strongest signal that practitioners were in the room. #### Zones and conduits when the plant predates the standard ISA/IEC 62443 gives the vocabulary this argument needs. A zone is a grouping of assets sharing a security requirement; a conduit is the permitted communication between zones; the discipline is that every flow crossing a boundary is a conduit somebody named, rather than a route that happens to work. The value of the standard in brownfield is not the target architecture, which most plants will never reach. It is that it forces the inbound pathways of the second section to be written down as objects with owners. What is achievable without stopping production. Drawing zone boundaries on paper against the plant as it actually runs, and enumerating every flow that crosses them, requires no outage and is usually the first time the true count of inbound paths exists in one place. Establishing a boundary at the OT and IT interface, and putting historian and reporting flows through it in the outbound direction, is achievable because those flows are already one-directional in intent even where they were built two-directional in fact. Retiring conduits nobody can name is achievable and is the cheapest security work available in an OT estate: the pathway opened for a commissioning contract that ended in 2019 is still there, and closing it costs a change window rather than a capital programme. What is not achievable, and should not be promised. Microsegmentation down to the individual controller on a flat, unauthenticated fieldbus, without an outage, is not a brownfield project. The protocols carry no identity, the devices cannot enforce policy, and the enforcement point therefore has to sit above the device, which is the compensating-control admission from the previous section restated as an architectural limit. Anybody selling per-device zero trust into a plant of that vintage is selling an enforcement point in front of the plant and describing it as an enforcement point inside it. The honest brownfield position is a small number of well-defended zones with a fully enumerated and mostly outbound set of conduits, achieved this year, in preference to a microsegmentation roadmap that is still a diagram in three years. #### The first ninety days The sequence matters more than any individual item in it, because the common failure is not omission but ordering: segmentation designed before the inventory exists, and microsegmentation scoped before anyone has counted the doors that open inward. Days one to thirty, find out what you have. Passive asset discovery, on the explicit understanding that active scanning can take a legacy device down. In parallel, and this is the part usually skipped, enumerate every inbound pathway: every forwarded port, every vendor tunnel, every dial-in, every VPN account, every remote-access tool installed for a project. Produce one list with a named owner and a live business justification against each entry. The firewall ruleset is not this list. It records what was permitted, not what is currently reachable, and the gap between the two is the finding. Days thirty-one to sixty, close and consolidate. Every pathway on the list without a current owner and a current justification is retired, and that will typically be a meaningful fraction of the total. Whatever survives is converted, in preference order: to outbound-initiated access where the vendor and the equipment allow it; failing that, into a single hardened, monitored, multifactor entry point rather than several unmonitored ones. Write down which pathways could not be converted and why, because that document is the honest description of your remaining exposure and it is what a regulator or an insurer will actually ask for. Days sixty-one to ninety, make it hold. Zone boundaries on paper against the plant as it runs, conduits named, and the segmentation policy managed out of band so that changes to it are visible. Monitoring concentrated at the boundaries, since the guidance is candid that OT environments "often lack robust logging, monitoring, and detection capabilities" and boundary visibility is what you can actually get. Break-glass access defined, time-limited and audited, before somebody needs it at three in the morning. Then rehearse the case where the pathway you rely on is the pathway that is compromised, because that is the scenario the architecture is for. #### What this does not solve Removing the listening surface removes one class of attack: the one that begins with an adversary finding you. It does nothing about the classes that begin elsewhere, and a document that stopped before saying so would be advocacy rather than doctrine. It does not address the insider, or the engineer with entirely legitimate credentials and poor judgement, because an architecture built on authenticated identity does exactly what an authenticated identity asks of it. It does not address the supply chain: firmware, an integrator’s laptop, a vendor’s update channel, or a compromise inside the software that establishes the outbound session itself. It does not address removable media, which is why the joint guidance devotes attention to media protection precisely because isolated plants attract USB as the transfer mechanism of last resort. It does not address availability and safety, which outrank it: any control capable of preventing an operator from reaching a plant during an incident will be bypassed in that incident, and should be. There is one limit specific to this architecture and it belongs in the open. If the protected machine reaches outward to a broker, then the broker, its operator and its software supply chain are in the trust path. The exposure has not been eliminated, it has been relocated, from an appliance on the plant boundary that anybody can scan to a component you can choose, place, monitor and, if necessary, replace. That is a considerably better trade than the one most plants have today, and it is a trade rather than a solution. Whoever operates that component should expect to be asked how it is built, where it runs, and what happens when it fails. Those are fair questions and this practice’s answer to them is a legitimate part of the assessment. Two controls follow from that admission and were missing from earlier revisions, which is an omission rather than a matter of emphasis: without them the architecture solves the old problem and creates a new one. The first is egress governance. Outbound-only is safe only where egress is constrained: destinations permitted by identity or name rather than left open, authenticated proxying, restricted name resolution, alerting on destinations never seen before, a standing prohibition on arbitrary tunnels, and certificate rotation and revocation defined before they are needed. A protected machine permitted to open any outbound connection it likes has been given a durable command-and-control and data-exfiltration path, engineered and signed off. The direction argument cuts both ways, and this is the edge that points inward. The second is a stated failure and compromise model for the broker. Whoever specifies this architecture should be able to say, before it is built, whether the broker is vendor-operated, customer-operated, single-tenant or shared; what fails open and what fails closed; which route is the safety-approved break-glass when the broker is the thing that is broken; how a session is cut mid-flight; which logs are retained and where they are held; how a connector identity is revoked; and how the plant keeps running safely when name resolution, internet transit, the broker or the identity provider is unavailable. The NCSC makes the same point for ordinary zero trust deployments, emphasising centralised logging and tightly governed, tested break-glass mechanisms. In operational technology it is not a refinement. A plant whose remote access depends on a component it does not operate has acquired a new single point of failure, and the only acceptable version of that is one where somebody has written down what happens when it fails and has rehearsed it at three in the morning. Sources: - Adapting Zero Trust Principles to Operational Technology, joint guidance, 29 April 2026: https://www.ic3.gov/CSA/2026/260429.pdf - Zero trust network access reference architectures, NCSC, which describe the outbound reverse-tunnel pattern this document had characterised as its own: https://www.ncsc.gov.uk/collection/zero-trust/zero-trust-network-access-ztna/ztna-reference-architectures - NIST SP 800-207A, A Zero Trust Architecture Model for Access Control in Cloud-Native Applications in Multi-Cloud Environments, September 2023: https://csrc.nist.gov/pubs/sp/800/207/a/final - NIST SP 800-207, Zero Trust Architecture: https://csrc.nist.gov/pubs/sp/800/207/final - NIST SP 800-82 Revision 3, Guide to Operational Technology Security: https://csrc.nist.gov/pubs/sp/800/82/r3/final - ISA/IEC 62443, zones and conduits: https://www.isa.org/standards-and-publications/isa-standards/isa-iec-62443-series-of-standards ### Governing agentic systems Version 1.2, revised 2026-07-27. An autonomous agent holding credentials is an access problem before it is a model problem. None of the major frameworks was written for this. Plain summary: Organisations are starting to use software that acts on its own. It can log in, make choices, and carry out tasks without a person approving each step. These are called agentic systems. The main rulebooks for managing artificial intelligence were written before this became common. They deal with how a model is built and tested. They say much less about software that holds a password and acts. This document argues that such a system should be governed the way a member of staff is governed: given only the access it needs, watched while it works, and able to be stopped. It also argues that this is why many of these projects stall. They are built around what the software can be shown to do, rather than around a specific business problem and the way the organisation would actually run it. #### The gap in the frameworks The frameworks an institution will be asked about were written to govern a model: how it was trained, what data went into it, how it was evaluated, what it is permitted to be used for. NIST's AI Risk Management Framework organises that work into Govern, Map, Measure and Manage. ISO/IEC 42001 makes it a certifiable management system. Regulation (EU) 2024/1689 attaches obligations to high-risk uses and to the providers and deployers behind them. All three are competent at what they set out to do, and none of them was written for software that authenticates to your systems and then acts without a person approving each step. The most precise way to show the gap is to look at what the EU AI Act assumes. Article 9 requires a risk management system running across the lifecycle. Article 14 requires that a high-risk system be capable of effective human oversight, including the ability "to intervene in the operation of the high-risk AI system or interrupt the system through a stop button or a similar procedure that allows the system to come to a halt in a safe state". Article 43 provides for conformity assessment before the system goes to market. Read those together and a shared premise appears: that the relevant behaviour of the system is knowable, and stable, at the point of assessment. That premise is what an agent breaks. Its behaviour at run time is a function of the tools it can reach, the content it encounters and the sub-tasks it chooses, and the interesting question is not what it does but what it is able to do. Notably, the Regulation does not define "agentic systems" at all. The obligations are not wrong; they are attached to the wrong noun. This is no longer a reading that has to be argued against institutional silence, because the institutions have conceded it. On 8 January 2026 NIST's Center for AI Standards and Innovation opened a request for information, docket NIST-2025-0035, asking specifically about gaps in existing cybersecurity frameworks when they are applied to AI agents. A standards body does not ask that question about a solved problem. On 17 February 2026 NIST announced an AI Agent Standards Initiative, with control overlays for SP 800-53 among the intended outputs and still in development. The gap is now a matter of record. One jurisdiction has already published for agents specifically, and it is a regional one. Singapore's Infocomm Media Development Authority issued the Model AI Governance Framework for Agentic AI on 22 January 2026 and updated it on 20 May 2026, adding case studies from real deployments and new material on multi-agent systems, third-party agents and automation bias. Its accountability position is the single most useful line an adviser can take to a risk committee. While agents may act autonomously, human responsibility continues to apply, and organisations and humans remain accountable, as deployers, for the decisions and actions of agents. That is close to this document's own thesis and arrives at it from the other direction. IMDA reasons from accountability to controls; the argument here reasons from the controls back to accountability. Where this document goes further is in naming which discipline owns the problem, and the answer is not the model governance function. #### The same conclusion, reached from engineering A reasonable objection to everything above is that it is a security argument, made by someone whose discipline predisposes them to it. The useful reply is not a louder security argument. It is that the same conclusion has been reached independently, from engineering, by people who were not thinking about governance at all. Ericsson's paper on AI agents in telecom network architecture is written for network architects and reasons from 3GPP network functions, the TM Forum intent management framework and the O-RAN service management layer. Its summary position is that "we should consider AI and agents as realization techniques for functions in the network architecture such as IMFs, rApps, or 3GPP network functions, not as a standalone function or an architecture". Strip the telecom vocabulary and that is the claim this document has been making: an agent is not a new kind of thing needing a new kind of governance, it is a way of implementing a function that already sits inside an authorisation structure. Two disciplines, two vocabularies, no shared regulator, same answer. The paper also supplies the plainest available test for where autonomy actually begins, which is a question institutions get wrong in both directions. A restricted agent becomes unrestricted, it says, in two cases: "Modification of internal logic: Overriding human-programmed restrictions" and "Modification of goals: Lifting the boundaries of human-assigned goals". That is a far more useful boundary than the usual one, because most of what is sold as an autonomous agent modifies neither. It selects among tools it was given, in pursuit of a goal it was set. It is worth knowing which of those you have bought. It is also worth borrowing the paper's treatment of the copilot, because it fills a real gap in the five questions below. A copilot is defined there as "a restricted, LLM-based agent designed to work interactively with humans as a human-to-machine interface". Naming it as a subtype rather than a separate category settles a scope argument that otherwise recurs at every risk committee: a copilot is in scope for the register and for the authorisation record, and its human-in-the-loop design is a control on the second question rather than an exemption from the first. One distinction has to be stated explicitly, or the borrowing does more harm than good. Classifying an agent as restricted is a description, not a control. Nothing about the label prevents the system from modifying its own goals; it records that it is not supposed to. What detects a restricted agent behaving like an unrestricted one, and what intervenes when it does, is the apparatus this document has spent five sections on: the per-instance identity, the deterministic limit, the append-only trail and the tested stop. A taxonomy tells you what you believe you have. The controls tell you whether you still have it. #### An agent is an identity, not a model The moment a system authenticates and acts, the governing questions stop being questions about a model and become the four questions any security architect asks about any principal on a network. What can it reach. Under whose authority. How do we know it was this one. How do we take that away. Accuracy, bias and evaluation remain necessary and stop being sufficient, because none of them tells you the blast radius of a system that holds a credential. A model that is wrong produces a bad answer. An agent that is wrong, or is steered, executes. The consequence is a change of ownership, and it is the practical reason this matters more than it sounds. In most institutions model risk sits with a data science or model risk function, and identity, entitlement and revocation sit with security and with IAM. An agent is the first artefact that belongs to both and is usually governed by neither, because each function can see only its half. The remedy is not a new committee. It is to enter the agent in the register where principals are already recorded, and to make it subject to the controls that already govern principals. The institution best placed to contradict this has instead reached for the same toolkit. NIST's National Cybersecurity Center of Excellence published a concept paper on 5 February 2026 on accelerating the adoption of software and AI agent identity and authorisation, and the standards it proposes to build on are OAuth 2.0, SP 800-207 zero trust architecture, and SP 800-63-4 digital identity. That is an identity and access problem being addressed with identity and access machinery. Note carefully what it is: a concept paper describing a potential collaborative project, with its comment period closed in April 2026. It is not guidance, and that is the more useful fact rather than a weakness in the citation. The problem has been named by the standards body and solved by nobody, which is precisely the space an institution deploying agents this year has to act in. This is also the seam between the two documents on this site, and the reason they belong together rather than being two subjects one adviser happens to cover. The first argues that access should be arranged so that nothing is left listening. The second argues that an agent should be scoped like a principal. They are the same architecture applied twice: authenticated identity at both ends, authorisation as a property of the peer rather than of a network position, and no standing exposure between engagements. An institution that has done the first work has most of the machinery for the second. Nothing in any current framework says this, which is the argument for writing it down. #### What least privilege means for something that reasons Least privilege was formulated for actors whose repertoire is enumerable. A batch job does the same thing tonight as it did last night; a member of staff has a role with a description. An agent composes: it selects tools, orders them, and can reach a state nobody in the design review considered, without any individual permission having been exceeded. This is the practical difficulty, and it is why "give it only what it needs" is true and insufficient as an instruction. What transfers from existing zero trust thinking is most of the mechanism. Per-request authorisation rather than a session that is trusted once. Short-lived credentials. Authorisation evaluated against the peer and the context rather than the network location. Explicit deny by default. All of it applies to an agent unchanged, and an institution that has built it for humans and workloads does not need a new stack. What does not transfer is the assumption that the union of permitted actions is safe because each action is permitted. Read the tool inventory as a set and ask what the combination makes possible: read a customer record, draft a message, send it externally; query a position, calculate, submit an instruction. Each is legitimate. The chain is the exposure, and it is invisible to a permission-by-permission review. This is why the OWASP Top 10 for Agentic Applications separates ASI02, tool misuse and exploitation, from ASI03, identity and privilege abuse: the first is the abuse of things the agent was legitimately given, the second is holding more than it should. Most governance effort goes into the second, and the first is the one that survives a tidy entitlement review. Two limits, stated in the language a risk committee uses. The scope of an agent must never exceed the scope of the human or department authorising it, which the IMDA framework states directly: authorisation should be tied to a supervising agent, a human user or an organisational department, bounded by session or time, non-transferable, and no wider than the permissions of the authorising human. And the limits should be deterministic. IMDA's wording is that limits should be preferred deterministic rather than non-deterministic, and bound by design. Put plainly: a boundary enforced by the instruction not to cross it is not a boundary, because the thing being instructed is the thing being constrained. Approval gates for irreversible actions, hard ceilings on value and volume, and refusal by default when the approval path is unavailable are limits. A prompt is a preference. #### Attribution and the audit trail When an agent acts, who acted? A regulated institution needs an answer that survives an examination, and the honest position is that the usual logging estate does not produce one. Conventional monitoring records the terminal event: the API call, the transaction, the record that changed. For a human that is close to sufficient, because a person can be asked what they were doing. For an agent the terminal event is the end of a chain, and the chain is the part under question. The difficulty is specific and worth stating precisely rather than as a general complaint about visibility. Inside a single task an agent may call many tools, revise its plan on what they return, and spawn sub-agents that call further tools under a delegated identity. The intermediate reasoning that selected those actions is inside the model. What the SIEM receives is the last hop, correctly authenticated and entirely unexplained. Three questions follow that most institutions currently cannot answer: which agent instance acted, under whose authority, and what caused it to choose this action rather than another. What the answer has to contain is now reasonably clear from the published material, and it is more than log retention. The IMDA framework asks for records of agent actions, decisions and interactions across all components, for the agent's plan and reasoning to be logged so it can be evaluated and verified, and for log immutability such that problematic agent trajectories and failures cannot be deleted, preserving them for analysis and compliance. A tamper-proof trail is also the evidence chain for a post-incident investigation. Translated into design terms, that is four requirements. The identity has to be per-instance rather than a shared service account. Delegation has to be carried in the trail rather than inferred from it. The plan has to be recorded as it is formed rather than reconstructed afterwards. And the record has to be append-only, in a place the agent cannot reach. The last one is the one most often missed, and it is the one that matters most: an agent with write access to the store holding its own trail has, by construction, the ability to edit the evidence. A note on evidence. There are widely circulated survey figures about how few organisations monitor agents end to end. Most originate with vendors selling the remedy, and this practice does not cite a number it cannot trace to a stated method and date. So none is quoted here. The four requirements above are drawn from published framework text rather than from a market statistic, which is the more durable basis for them in any case: a figure about how many institutions currently fall short would date within a year, while what an audit trail has to contain does not. #### The stop button The obligation exists in law already. Article 14 of the AI Act requires that oversight staff be enabled to intervene in the operation of a high-risk system or interrupt it through a stop button or similar procedure "that allows the system to come to a halt in a safe state". Note the last clause, because it does most of the work: stopping is not enough, the state you stop in has to be safe. Anyone who has worked in operational technology will recognise the requirement; it is the same reason an emergency stop on a production line is engineered rather than merely wired. For an agent, halting in a safe state has to be designed, and there are at least four things behind the button. Revocation that actually takes effect, which means short-lived credentials rather than a long-lived key you must now chase through every downstream system. Session termination that ends work in progress rather than letting a queued instruction land after the agent is gone. Reversal, or an explicit record that a given class of action is irreversible, which is the real reason payments, external communications and disclosures deserve approval gates rather than a faster stop. And containment of delegation, because stopping a supervising agent while its sub-agents continue under credentials it issued is not a stop. IMDA is direct about the surrounding behaviour, in three parts. Mechanisms and procedures should be designed to take agents offline and limit their potential scope of impact when they malfunction. Action should be denied by default when approval infrastructures fail. And in the event of catastrophic malfunction or compromise, commensurate measures such as termination and fallback solutions should be considered. The middle one is the one to argue for hardest, because it is where an institution learns whether it has a stop button or a preference. If the approval path is unavailable, say the supervisor cannot be reached, the safe default is refusal. Systems built for availability will proceed. That default is a decision somebody should make deliberately, in advance, in writing. The test is procedural, not technical, and it is the one question in this document an institution can answer this week. Who is authorised to stop an agent in production, at three in the morning, without needing anyone else's approval, and when did they last do it as a drill? An untested stop is a claim. It also has to be operable by whoever is on shift, which means it cannot live only with the team that built the agent. #### Why these programmes stall Everything above concerns what has to be true before an agent is trusted with anything. In practice those requirements are met late, when the thing is already built and someone asks whether it can go live. That sequence is itself the finding. Most initiatives that stall do not stall because the model underperformed. They stall because they were designed around what the model could be shown to do, rather than around a bounded business problem and the operating model required to solve it. The distinction matters because the two produce different artefacts. A demonstration is optimised for a room: a plausible task, a cooperative example, a narrator to supply the context. Everything genuinely difficult about the production version is supplied by the person presenting, and is therefore invisible in the thing being judged. What is missing is not capability. It is the operating model: who the work belongs to, what the agent is permitted to decide alone, where the institutional knowledge comes from that the demonstration supplied by hand, which systems it may read and write, what stops it, who it escalates to, and what would count as success in numbers somebody already reports. Those seven elements, workflow, decision authority, enterprise knowledge, data access, controls, human escalation and measurable outcomes, are not a sequence of gates. They are constraints on one another, and designing them together is the whole of the work. Decision authority is meaningless without an escalation path that a named human answers. Data access cannot be scoped until the workflow is known. Measurable outcomes cannot be defined without a baseline, which usually has to be measured before anything is built. An initiative that settles six and defers one does not arrive six-sevenths of the way; it arrives at a pilot that cannot be promoted, which is the state most of them are in. Gartner forecasts that more than forty per cent of agentic AI projects will be cancelled by the end of 2027, and attributes it to escalating costs, unclear business value and inadequate risk controls. Two cautions about that figure, both of which matter more than the number. It is a forecast rather than a measurement, so it evidences an analyst view and not an outcome. And a prediction of cancellation is not a prediction of technical failure. Read the three causes as a description of what was never designed: cost that was not bounded because the workflow was not, value that is unclear because no outcome was defined before building, and controls that are inadequate because they were scoped after the fact. The same source notes that where an agent is genuinely warranted, rethinking the workflow from the ground up is usually the path that works, which is the same claim from the other direction. One consequence is worth stating plainly for anyone sponsoring this work. The right first question is not which process could be given to an agent. It is which decision is currently slow, expensive or inconsistent, what it would be worth to change that, and whether an agent is the cheapest way. Many use cases presented as agentic do not require an agent at all: a scripted automation, a better form or a retrieval tool would do the same work with a fraction of the governance burden. Establishing that early is not scepticism about the technology. It is what makes the cases that do warrant an agent defensible when they are challenged, and they will be challenged by someone holding this list. #### A minimum viable governance position The timing argument has to be made carefully now, because the deadline it used to rest on has moved. Until July 2026 this section argued that the EU enforcement date arrived before the guidance did. Regulation (EU) 2026/1744, the Digital Omnibus on AI, was published in the Official Journal on 24 July 2026 and entered into force on 27 July, deferring the Chapter III high-risk obligations for stand-alone Annex III systems from 2 August 2026 to 2 December 2027, and for systems embedded in regulated products under Annex I to 2 August 2028. Anyone still citing August 2026 for human oversight of a high-risk system is citing a repealed date, and this document did so until this revision. What did not move is worth more than what did. The Article 50 transparency obligations still apply from 2 August 2026: the deferral reaches Chapter III and not the rest of the Regulation. So an institution deploying agents that interact with people, or generate synthetic content, has an obligation this month rather than next year. And the deferral changes the deadline without changing the substance: conformity assessment, technical documentation, human oversight and registration all still arrive, for systems being built now and reviewed later against them. The underlying argument therefore survives the correction, and is better for losing the deadline as its crutch. NIST's substantive agent deliverables are still not in place: the COSAiS project lists five overlays, of which the predictive-AI one reached an annotated outline for comment in January 2026, and the two agent overlays, single and multi-agent, sit behind it in the sequence. So the gap between deploying agents and having a standard to deploy them against is real, measurable and unchanged by the Omnibus. What has changed is that it is now a gap of practice rather than a countdown to a penalty. An institution that waits for the standard still has no record of the period in which it was already deploying, and that record is the artefact a regulator asks for in December 2027 about the work done in 2026. What follows is the minimum an institution should be able to answer before an agent reaches production. It is deliberately five questions, because a governance position nobody can recite is a document rather than a control. One. What agents exist? A register, per instance and not per project, with each entry naming its purpose and its owner. IMDA calls the alternative agent sprawl and recommends a central catalogue to prevent it. This is the same claim that opens the third document on this site: inventory is the first control, and unowned means unmanaged. Two. What can each one reach, and who authorised that? The tool and data inventory per agent, read as a set rather than a list, with the authorising human or department named. No agent's scope should exceed the scope of the person or department authorising it. Three. Who owns it, and who is accountable when it is wrong? A named person, not a team. The agent is not a principal and cannot hold the accountability; a human does, and this is where the IMDA position becomes operational rather than declaratory. Four. How is it stopped, and when was that last tested? Named authority, out of hours, no second approval, with a date against the last drill. Untested means unproven. Five. What evidence survives the incident? Per-instance identity, the recorded plan, the delegation chain, and an append-only store the agent cannot write to. If a reconstruction of what the agent did depends on the agent's own logs, there is no evidence. Two reference points make the checklist defensible in front of people who will test it. The OWASP Top 10 for Agentic Applications for 2026 is the threat taxonomy to map against, and its ten categories, from goal hijack and tool misuse through identity and privilege abuse to rogue agents, are the failures these five questions are trying to make survivable. The NCCoE concept paper is the architectural sketch to align with, so that work done now is not orphaned when the standards arrive. Neither is a compliance regime, and neither should be presented as one. What the five answers produce is an artefact: an inventory, an authorisation record, a tested stop and an evidence chain, dated. That is what an institution will be asked to show for the period before the guidance existed, and building it now is considerably cheaper than reconstructing it later. Sources: - NIST AI Risk Management Framework 1.0: https://www.nist.gov/itl/ai-risk-management-framework - ISO/IEC 42001:2023, AI management systems: https://www.iso.org/standard/42001 - Regulation (EU) 2024/1689, the AI Act: https://eur-lex.europa.eu/eli/reg/2024/1689/oj - Model AI Governance Framework for Agentic AI, IMDA Singapore, published 22 January 2026, updated 20 May 2026: https://www.imda.gov.sg/-/media/imda/files/about/emerging-tech-and-research/artificial-intelligence/mgf-for-agentic-ai.pdf - Regulation (EU) 2026/1744 of 8 July 2026, the Digital Omnibus on AI, amending Regulation (EU) 2024/1689, in force 27 July 2026: https://eur-lex.europa.eu/eli/reg/2026/1744/oj/eng - SP 800-53 Control Overlays for Securing AI Systems (COSAiS), NIST, five planned overlays including single-agent and multi-agent: https://csrc.nist.gov/projects/cosais - AI agents in telecom network architecture, Ericsson, 17 October 2025, revised 21 July 2026: https://www.ericsson.com/en/reports-and-papers/white-papers/ai-agents-and-network-architecture - Request for Information: Security Considerations for AI Agents, NIST CAISI, docket NIST-2025-0035, 8 January 2026: https://www.federalregister.gov/documents/2026/01/08/2026-00206/request-for-information-regarding-security-considerations-for-artificial-intelligence-agents - AI Agent Standards Initiative, NIST, 17 February 2026: https://www.nist.gov/news-events/news/2026/02/announcing-ai-agent-standards-initiative-interoperable-and-secure - Accelerating the Adoption of Software and AI Agent Identity and Authorization, NIST NCCoE concept paper, 5 February 2026: https://csrc.nist.gov/pubs/other/2026/02/05/accelerating-the-adoption-of-software-and-ai-agent/ipd - OWASP Top 10 for Agentic Applications, 2026: https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ - Gartner forecast: over 40 per cent of agentic AI projects will be cancelled by the end of 2027, 25 June 2025: https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027 - NIST IR 8596, Cybersecurity Profile for Artificial Intelligence, initial public draft: https://csrc.nist.gov/pubs/ir/8596/iprd ### Cloud spend is attack surface Version 1.1, revised 2026-07-28. Unowned resources are unpatched resources. Cost management read as a control discipline rather than a procurement exercise. Plain summary: When an organisation moves to cloud computing, it becomes easy for teams to create new systems. That is mostly a good thing. The side effect is that the number of systems, and the question of who is answerable for each one, gets harder to answer over time. This is usually handled as a cost question. It is also a security question. A system without a clear owner tends not to be updated, and an out-of-date system is among the easiest ways in. This document argues that knowing what you run, and who owns it, is a security control as much as a budgeting exercise. #### Why cost lands on the wrong desk A cloud invoice is the most complete inventory of running systems that most institutions possess. It is itemised, it is produced monthly without anyone being asked, and unlike every asset register maintained by hand it cannot omit a resource, because the provider is charging for it. That document is delivered to finance and procurement, where it is read for a single purpose: is the number larger than last month, and can it be made smaller. The consequences of what is on that list land somewhere else. An instance nobody owns is not a budget variance, it is an unpatched host. A storage bucket created for a migration that finished two years ago is not a rounding error, it is data with no reviewer. The costs are read by one function and the risks are carried by another, and the two functions are looking at the same list. The mismatch is structural rather than anybody's failure. Finance has the document and no mandate to ask who owns line item four hundred and twelve. Security has the mandate and works from an asset register that was assembled by asking teams what they run, which is to say a register that is complete only where people remembered. Meanwhile the platform made creating resources a self-service action, correctly, because that is most of the value of moving. Nobody made forgetting them cost anything. The rest of this document is about reading the bill as the control artefact it already is. #### Unowned is unpatched The claim is narrow and mechanical. Patching, configuration review, log review and decommissioning are all activities that require somebody to decide to do them. When no owner exists, no such decision is taken, and the resource continues running in the configuration it had on the day it was created. Nothing has to go wrong for this to happen; it is the default state of an unowned system, and drift is the only thing operating. This is not a FinOps argument dressed up as a security one. Established security doctrine already puts inventory first. Inventory and Control of Enterprise Assets is CIS Critical Security Control 1, ahead of anything about patching, logging or monitoring. It asks organisations to actively manage all enterprise assets connected to the infrastructure "physically, virtually, remotely, and those within cloud environments, to accurately know the totality of assets that need to be monitored and protected". It also asks explicitly for the identification of "unauthorized and unmanaged assets to remove or remediate". Asset management sits in the Identify function of the NIST Cybersecurity Framework 2.0, upstream of Protect. The doctrine has always been that you cannot defend what you have not enumerated. So the argument in this document is not that people who manage cost should care about security. It is narrower and harder to dismiss: the enumeration that Control 1 asks for already exists, it is produced monthly by the cloud provider, and it is being read by the wrong function for the wrong purpose. Cost allocation and asset inventory are the same exercise conducted with different intent, and only one of the two has an unavoidable, complete, monthly data source. One boundary, because the claim is easy to overstate and this is where a sceptical reader should push. An owned resource is not automatically a patched resource. Ownership creates an accountable party; it does not perform maintenance, and a resource with a nominal owner who does nothing is no better patched than an orphan. What ownership changes is that the omission becomes attributable and can therefore be measured, reported and escalated. That is the entire claim. It is smaller than "FinOps improves security" and it is defensible, which is worth more. #### What the survey showed The survey behind this section was titled Malaysia's Cloud Survey 2020. The slide presenting it records that it was led by Alphaus Cloud in collaboration with Hitachi Sunway, and supported by the MDEC Cloud Department, MAJECA, MASSA, UUM, PIKOM and Integra. Mohd Atasha presented the findings at the inaugural KL Cloud Maestro Series on 2 October 2020. Two results have been carried on the public record since then and are repeated here unchanged: half of respondents cited missing skill sets as the largest gap after adoption, and four in five named scalability as the principal benefit they sought. Read together those two numbers describe a specific and predictable failure. Organisations adopted the cloud for scalability, which is a capability for creating resources quickly, while identifying their own largest post-adoption gap as skills. Rapid creation and thin operational capacity is the condition under which unowned resources accumulate, and it accumulates fastest in exactly the organisations that got the most value from adopting. This is not an argument against scalability. It is an argument that the skills gap the respondents named themselves has a security consequence they were not describing at the time, because it was being discussed as a cost and capability question. The method can now be stated, and it matters more than the percentages do. It was an online survey, undertaken because the industry data the organisers wanted did not exist: in the presenter's own words, "we don't get the data", so the survey was run rather than the gap assumed. There were 105 respondents. The average size of a participating organisation was about 1,200 people, and the respondents held senior technology and executive roles. The slide names a selection of the participating organisations, among them Hitachi, Etiqa, Celcom, TIME, Schlumberger, Boost, AT&T, Gamuda, TM and Sunway Medical Centre, which places the sample among large Malaysian enterprises rather than among small businesses. A stated n of 105 in that population is a modest but real sample, and it is worth considerably more than a percentage with no denominator behind it. The same slide carries three findings that are not the two quoted above and are stated here because they are legible on the record: 77 per cent of respondents had adopted cloud in some form, 45 per cent used public cloud and 37 per cent used a hybrid arrangement. Those figures describe the shape of adoption rather than its consequences, which is why the two results at the head of this section remain the ones the argument rests on. Three limits, because the point of publishing a method is that a reader can find its weaknesses. The duration is recorded inconsistently in the source itself: the slide states three weeks and the presenter says about two, and rather than choose the more convenient number, both are noted here. The question wording and the sampling frame, meaning how participants were approached and who could have answered but did not, are still not published, so this remains a survey of those who chose to respond. And the surviving recording is a short extract that shows only the summary slide, so it establishes the method above but does not itself display the skills and scalability figures; those rest on the presenter's own record of the survey rather than on the recording. Each of those is a reason to treat the two headline results as an indication rather than a measurement, and saying so is cheaper than being corrected later. #### Ownership as a control, not a spreadsheet column Tagging is usually implemented to answer the question "which cost centre does this belong to", and the answer is a code. It could as easily answer "who is accountable for the security state of this resource", and the answer would be a person. The instrumentation is identical; only the intent differs, and the intent determines whether the output is usable as a control. Four properties separate ownership that functions as a control from a spreadsheet column. It names a person or a defined role, not a cost centre, because a cost centre cannot patch anything. It is enforced at creation, so that an untagged resource is either refused or quarantined, since retrospective tagging campaigns are how organisations discover that nobody remembers. It is validated against a live source, because the most common failure is not the absent tag but the tag naming somebody who left in March, and an owner who no longer exists reads as ownership while providing none. And unallocated spend is treated as a finding with an assignee and a date rather than a residual percentage, which is the single highest-value change available: in most estates the unallocated line is where the orphans are, and it is currently a rounding note in a monthly report rather than a security queue. The reason to reuse the cost mechanism rather than build a parallel one is that the cost mechanism is the only one with an unavoidable and complete monthly refresh. A security asset register decays the moment people stop updating it. The bill does not decay, because the provider has a commercial interest in its completeness. Attaching accountability to that artefact means the control is refreshed by somebody else's billing system, which is a considerably more reliable arrangement than depending on institutional memory. #### Where the FinOps framework stops The FinOps Framework is a competent instrument for what it was built to do, and this document has been leaning on its machinery throughout. Its capabilities include allocation, tagging and showback, which are precisely the mechanisms the previous section reuses. It is worth being exact about where it stops, because overstating it would undermine the argument. It optimises cost. Its purpose is to make spend visible, attributable and defensible, and the questions it is designed to answer are whether a resource is worth what it costs and who should carry the charge. It does not ask whether a resource is patched, whether its configuration has been reviewed, or whether the person named against it is still employed. A mature FinOps practice will happily allocate one hundred per cent of the spend on an unpatched instance to the correct cost centre and report that as success, because by its own terms it is. The seam is therefore precise. FinOps produces the enumeration and the attribution; it does not produce the security decision that follows from them, and nothing in the framework requires anybody to make it. Closing that seam is an organisational act rather than a tooling one, and it is small: the allocation report has a second reader, unallocated spend has an assignee and a due date, and the ownership field is validated against the identity directory rather than a spreadsheet. None of that is in the framework. All of it is available to anyone already doing the framework. Two limits belong in the open. This discipline addresses the resources you are billed for, which excludes anything running in an unbilled account, on premises, or under a personal payment card, and shadow IT of that kind is exactly the case where the bill is not the inventory. And enumeration plus ownership is where security work starts rather than where it concludes: the vulnerability management, configuration and log review still have to happen. The claim here is only that they cannot happen on a resource nobody has enumerated, and that the enumeration already exists. Sources: - FinOps Foundation Framework: https://www.finops.org/framework/ - CIS Critical Security Control 1, Inventory and Control of Enterprise Assets: https://www.cisecurity.org/controls/inventory-and-control-of-enterprise-assets - NIST Cybersecurity Framework 2.0: Identify function: https://csrc.nist.gov/pubs/cswp/29/the-nist-cybersecurity-framework-csf-20/final - FinOps Framework capabilities, including allocation and tagging: https://www.finops.org/framework/capabilities/ - Are Malaysian companies ready for cloud adoption? Survey findings, KL Cloud Maestro Series, 2 October 2020: /notes/cloud-adoption-survey/ ### The recommendation nobody acted on Version 1.0, revised 2026-08-04. Every argument here ends in an ask that requires somebody to be told they were wrong. Treating that as a communication problem is why sound work stalls. Plain summary: The other documents on this site each end with something an organisation should do. None of those things is difficult to understand. They are difficult for a different reason. Acting on them usually means somebody has to say that a decision already made and already approved turned out to be wrong. In many organisations that is an unsafe thing to say. So the work stops, and it is recorded as a technical problem or a budget problem, because those are easier to write down. This document argues that the ability to report bad news is part of the security control, not a matter of workplace atmosphere. It is measurable, it is documented in safety research, and it is the thing to establish before the recommendation is written rather than after it is ignored. #### Three documents, one shared point of failure The three arguments published alongside this one were written independently, about different domains, for different audiences. They arrive at recommendations that share a structure, and the shared structure is the subject of this document. The operational technology argument asks an operator to accept that the inbound pathways accumulated over a decade were each approved by somebody, and that the total was never anybody’s document. The agentic systems argument asks an institution to adopt refusal by default when an approval path is unavailable, which means accepting that a system built for availability was built to the wrong preference. The cloud spend argument asks an organisation to attach an owner to every line on an invoice, which surfaces, item by item, how much of the estate nobody could account for. None of those three is intellectually difficult. Each of them requires a named person to say, in front of colleagues, that something they approved or ran or signed for is not what they believed it was. That is the actual work, and it is the part no architecture diagram contains. It would be convenient to treat this as a separate discipline, to be handed to a communications function once the technical position is settled. The argument here is that this is precisely the wrong sequence, and that the conditions under which a recommendation can be acted on are a design constraint on the recommendation itself. #### Information flow is a safety property, and it is measurable The sociologist Ron Westrum published a typology of organisational cultures in a patient safety supplement of Quality and Safety in Health Care in 2004. His claim is narrow and useful: because information flow both influences performance and indicates other aspects of culture, it can be used to predict how an organisation will behave when signs of trouble arise. He sorted organisations into three kinds, distinguished by what happens to the person carrying the bad news. In a pathological organisation, messengers are punished. In a bureaucratic one, messengers are neglected. In a generative one, messengers are trained. The row worth sitting with is the one about failure. Under a pathological culture, failure leads to scapegoating. Under a bureaucratic culture, failure leads to justice, which sounds acceptable until you notice that justice means establishing who was at fault. Only in the third case does failure lead to inquiry. An organisation can be scrupulously fair, follow every procedure, and still be one in which nobody volunteers a problem early, because the reliable consequence of raising one is an investigation into who caused it. This matters here for a reason specific to the subject matter. Westrum was writing about safety, in aviation and healthcare, where the cost of an unreported anomaly is measured in lives. The finding was later carried into technology by the DORA research programme, which reports that a high-trust, generative culture of the kind Westrum describes predicts software delivery performance and organisational performance. So the claim is not that pleasant workplaces are more productive. It is that the handling of bad news is a property of a safety-critical system, established in the literature of safety-critical systems. It is also measurable, which is what separates this from an appeal to values. The Westrum construct is operationalised as six statements, scored on agreement: that information is actively sought; that messengers are not punished when they deliver news of failures or other bad news; that responsibilities are shared; that cross-functional collaboration is encouraged and rewarded; that failures are treated primarily as opportunities to improve the system; and that new ideas are welcomed. Six questions, a numeric baseline, and a defensible way to say whether the condition improved. An institution that will not measure this is not declining a soft initiative, it is declining to instrument a variable its own incident history depends on. #### The blameless retrospective is security guidance, not management theory A reader in a security function is entitled to be suspicious of an argument about behaviour, because a great deal of what arrives under that heading is unfalsifiable. So it is worth noting who else is making it. In October 2024, the Cybersecurity and Infrastructure Security Agency, the Federal Bureau of Investigation and the Australian Signals Directorate’s Australian Cyber Security Centre published a joint guide on safe software deployment. Its conclusion names two approaches for keeping a deployment process inside its safety boundary. The first is to “foster a blameless retrospective (also called postmortem) culture, where teams analyze both positive and negative outcomes by focusing on the processes that contributed to the result, rather than assigning blame to any individual”. It then states the design principle underneath that, in one sentence: “Individual actions should not lead to an incident if the environment and processes are resilient.” That sentence is doing more work than it appears to. It relocates the question. If an individual action was sufficient to cause an incident, the finding is about the environment that permitted it, and an investigation that terminates at the individual has stopped one step early. This is a technical claim about system design, issued by three national security agencies, and it is the same claim Westrum arrived at from accident research twenty years earlier. The same guide asks for two further things that are behavioural rather than technical, and are easy to skim past. Organisations should “establish a culture of encouraging staff to report potential problems, even when the problems seem negligible”. And near misses should be treated as real incidents, because they “provide an opportunity to enhance the program without the software manufacturer or their customers experiencing the full negative impact of an actual incident”. A near miss is only available to an organisation whose staff report things that did not go wrong. That reporting is the control. Nothing else in the deployment pipeline can substitute for it. The underlying construct has a longer empirical history. Amy Edmondson’s 1999 study in Administrative Science Quarterly introduced team psychological safety and tested it in a field study, finding it associated with learning behaviour, of which reporting error is the clearest instance. It is among the most heavily cited papers in organisational research. The relevant point for this argument is modest and specific: the willingness to say that something is wrong varies measurably between teams doing the same work in the same institution, and it is a property of the team’s conditions rather than of the character of its members. #### What his own survey found, read again There is one piece of evidence in this argument that is not drawn from published literature. In 2020 Mohd Atasha presented the findings of Malaysia’s Cloud Survey 2020 at the inaugural KL Cloud Maestro Series, on 2 October. The survey was led by Alphaus Cloud in collaboration with Hitachi Sunway and supported by the MDEC Cloud Department, MAJECA, MASSA, UUM, PIKOM and Integra. It was an online survey with 105 respondents in senior technology and executive roles, the average participating organisation being around 1,200 people. The survey was run because the industry data the organisers wanted did not exist. Two findings have been on the public record since. Half of respondents named missing skill sets as the largest gap after adoption. Four in five named scalability as the principal benefit they were seeking. Those two answers came from the same people about the same programmes, and read together they describe something other than a technology problem. The capability being bought was the ability to create infrastructure quickly. The constraint being reported, by the buyers themselves, was human capacity. Nobody in that sample said the technology did not work. They said their organisations could not absorb it at the rate they were adopting it. The FinOps document on this site reads the same finding as a security argument, and it holds: thin operational capacity plus rapid creation is the condition under which unowned resources accumulate. Read for this document, it says something adjacent. When practitioners were asked what was hardest about a major technology change, the answer they gave was about people, and it was the largest single gap they identified. That was 2020, about cloud adoption. The recommendations in the other three documents here make heavier demands of an organisation than a migration does. The limits of this evidence should be stated. The surviving recording is a short extract showing the summary slide, so it does not display the skills and scalability figures, which rest on his own record of the presentation. The slide records a three-week fielding period where the presenter says about two. The exact question wording is not published. Those are the same caveats carried in the FinOps document, and they are repeated rather than dropped, because a number reused in a second argument is not thereby better attested. #### Designing the ask so it can be accepted The literature offers stage models, most famously an eight-step sequence, and they are usually introduced with the claim that seventy per cent of change initiatives fail. That statistic should be treated with care. Mark Hughes examined its provenance in the Journal of Change Management in 2011 and found the figure repeatedly asserted and attributed without traceable empirical support. This document declines to use it, and declines the stage models with it, on the grounds that a site inviting readers to check its claims should not build on a number that cannot be checked. What can be offered is narrower: a set of properties that make a recommendation acceptable to the organisation that has to act on it. Each follows from something already established above. Establish the reporting condition before the assessment, not after. If the Westrum items score badly, the assessment will return a shorter list of problems than the estate contains, because the people who know about them have accurately judged the consequence of saying so. An assessment conducted in that condition is not a measurement of the estate, it is a measurement of what is currently sayable. Six questions, asked first, tell you which of the two you are about to buy. Separate the finding from the person. The joint guidance already states the principle: individual actions should not lead to an incident if the environment and processes are resilient. Applied to an inbound pathway approved in 2014, the finding is that change control assessed whether access was justified and nothing assessed what remained listening afterwards. That is a structural gap with an owner who can fix it. The same fact expressed as a person’s error is unactionable, because the only available remedy is a reprimand and the pathway stays where it is. Make the first admission at the top. The demand in each of these documents is that somebody concede an error. The cheapest way to establish that this is survivable is for the most senior person present to do it first, about their own decision. This is the least technical item on the list and it is the one that most reliably determines whether the rest proceeds. It cannot be delegated downward, because its entire content is who was willing to go first. Ask for the near miss, and treat it as a finding. Near misses are the only cheap information an organisation gets, and they exist solely where people report things that did not go wrong. An institution that can produce a list of near misses has already demonstrated the reporting condition, and one that cannot has answered the question whether it is present. State the limits of the recommendation in the recommendation. Every document on this site carries a section on what it does not solve. That is partly honesty and partly a practical device: an argument that names its own boundary invites disagreement about the boundary rather than about the author, and disagreement about the boundary is a conversation that can end in a decision. #### What this does not solve This argument does not work where the information is unwanted. There are institutions in which the absence of reported problems is the desired state, because a reported problem creates an obligation to spend. Nothing in this document changes that, and an adviser who claims to be able to should not be believed. What can be done is to name the condition accurately and let the sponsor decide, rather than delivering an assessment whose optimism is an artefact of what nobody would say. The evidence is associative and its direction is not fully established. Westrum was careful about this himself, writing that the relationship between culture and safety “requires more exploration before the connection can be considered definitive”. It remains plausible that organisations performing well can afford to handle bad news generously, rather than that handling bad news well causes the performance. The practical consequence is a limit on what may be promised: improving the reporting condition is not a lever that produces a predictable delivery outcome, and the six items are a diagnostic rather than a dial. None of this substitutes for the technical work. A generative culture with no asset inventory still has no asset inventory. The claim here is only that the three technical arguments on this site have a common precondition for being acted on, and that the precondition is usually left unexamined while the technical content is debated. The precondition is not the work. It is what determines whether the work gets done. Finally, what this argument rests on, and what it does not. It rests on the published literature cited above and on one piece of first-party evidence, the 2020 survey read again in the section before this one, whose limits are stated where it is used. What it does not contain is a worked case: an account of one specific programme that stalled for this reason, named and traced from the recommendation to the point at which it stopped. That absence is a real limit and it is worth being precise about which kind. It is not a gap in the argument, which stands on the safety and security literature and does not depend on any single instance. It is a gap in the persuasion, because a reader who has watched this happen inside their own organisation will recognise the pattern immediately, and a reader who has not is being asked to accept it on the strength of the citations alone. The reason no case appears here is worth stating plainly rather than leaving as an omission. The programmes this pattern was observed in belong to the organisations that ran them, and the useful detail is exactly the detail that would identify the people who could not safely report a problem. An argument about the cost of unsafe disclosure would be a poor place to demonstrate indifference to it. A sufficiently anonymised case would be safe to publish and would also be unfalsifiable, which is the standard this site refuses elsewhere: a reader could not check it, and an uncheckable anecdote is weaker evidence than an honest statement that the case is absent. If an organisation is ever willing to be named, the case will be added and this paragraph will be replaced. Sources: - Ron Westrum, A typology of organisational cultures, Quality and Safety in Health Care 13 (2004), doi:10.1136/qshc.2003.009522: https://qualitysafety.bmj.com/content/13/suppl_2/ii22 - Safe Software Deployment: How Software Manufacturers Can Ensure Reliability for Customers, CISA, FBI and ASD ACSC joint guide, October 2024: https://www.ic3.gov/CSA/2024/241024.pdf - Amy C. Edmondson, Psychological Safety and Learning Behavior in Work Teams, Administrative Science Quarterly 44 (1999): https://doi.org/10.2307/2666999 - DORA, Generative organizational culture, including the six Westrum survey measures: https://dora.dev/capabilities/generative-organizational-culture/ - Mark Hughes, Do 70 Per Cent of All Organizational Change Initiatives Really Fail?, Journal of Change Management 11 (2011): https://doi.org/10.1080/14697017.2011.630506 - Are Malaysian companies ready for cloud adoption? Survey findings, KL Cloud Maestro Series, 2 October 2020: /notes/cloud-adoption-survey/ ## Frameworks ### Operational technology and network architecture - ISA/IEC 62443: Zones and conduits; industrial automation and control systems security. (https://www.isa.org/standards-and-publications/isa-standards/isa-iec-62443-series-of-standards) - NIST SP 800-82 Revision 3: Guide to operational technology security. (https://csrc.nist.gov/pubs/sp/800/82/r3/final) - NIST SP 800-207: Zero trust architecture. (https://csrc.nist.gov/pubs/sp/800/207/final) - Adapting Zero Trust Principles to Operational Technology: Joint guidance, April 2026. (https://www.ic3.gov/CSA/2026/260429.pdf) - NIST Cybersecurity Framework 2.0 (https://www.nist.gov/cyberframework) - MITRE ATT&CK for ICS (https://attack.mitre.org/matrices/ics/) - ISO/IEC 27001 and 27002 (https://www.iso.org/standard/27001) ### Artificial intelligence - NIST AI Risk Management Framework 1.0 (https://www.nist.gov/itl/ai-risk-management-framework) - ISO/IEC 42001:2023: AI management systems. Certifiable, and increasingly requested in procurement due diligence. (https://www.iso.org/standard/42001) - Regulation (EU) 2024/1689, the AI Act: Applies extraterritorially where outputs are used in the Union. (https://eur-lex.europa.eu/eli/reg/2024/1689/oj) - OECD AI Principles (https://oecd.ai/en/ai-principles) - Artificial Intelligence Systems Cyber Security Framework (AISCF): Malaysia, National Cyber Security Agency, launched 9 July 2026. Secures AI data, AI models, and AI infrastructure and applications in layers across a seven-phase lifecycle. Addressed to any organisation that builds, supplies or uses AI systems. He contributed to the taskforce and is acknowledged on page 58. The full document is free to download from NACSA. (https://nacsa.gov.my/aiscf.php) - Gartner: over 40 per cent of agentic AI projects will be cancelled by end of 2027: A forecast rather than a measurement, listed for the three causes it names: escalating cost, unclear business value and inadequate risk controls. (https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027) ### Financial services and operational resilience - Bank Negara Malaysia, Risk Management in Technology (RMiT): Policy document revised November 2025, effective 28 November 2025. The revision extends board accountability and adds explicit expectations on the governance of emerging technologies. (https://www.bnm.gov.my/documents/20124/938039/pd-rmit-nov25.pdf) - Monetary Authority of Singapore, Technology Risk Management Guidelines: Revised January 2021, with strengthened requirements on board oversight, secure development, emerging technology risk and cyber resilience. (https://www.mas.gov.sg/regulation/guidelines/technology-risk-management-guidelines) - APRA Prudential Standard CPS 234, Information Security (https://www.apra.gov.au/standards/cps-234) - Digital Operational Resilience Act (https://eur-lex.europa.eu/eli/reg/2022/2554/oj) ### Cloud financial management - FinOps Foundation Framework (https://www.finops.org/framework/) ## Credentials - AWS Certified Cloud Practitioner, Amazon Web Services: verifiable badge record (https://www.credly.com/badges/7f32a747-3a41-46a7-b8dc-adf87bb3fb97/public_url) - AWS Certified AI Practitioner (Early Adopter), Amazon Web Services: verifiable badge record (https://www.credly.com/badges/4096af86-8b64-4c3b-ad24-46c17db784ad) - FinOps Certified Professional, The Linux Foundation: verifiable badge record (https://www.credly.com/badges/aeff4ee5-5a4d-4711-b604-2a15052594ae/public_url) - FinOps Certified Instructor, The Linux Foundation: verifiable badge record (https://www.credly.com/badges/09633cf3-77b6-4938-9d0e-cedd835b6d66/public_url) - HRD Corp Accredited Trainer, Human Resource Development Corporation, Malaysia: verifiable trainer record, accredited 25 July 2024 to 25 July 2027 (https://trainers.hrdcorp.gov.my/ecert/d9950320-21b3-11ef-8688-c172cce5a103/d9950320-21b3-11ef-8688-c172cce5a103) - Certificate of Competence in Zero Trust (CCZT), Cloud Security Alliance: badge record to be added (https://cloudsecurityalliance.org/education/cczt) - FinOps Certified Engineer, FinOps Foundation (https://learn.finops.org/path/finops-certified-engineer) - PyTorch and Deep Learning for Decision Makers (LFS116), Linux Foundation (https://training.linuxfoundation.org/training/pytorch-and-deep-learning-for-decision-makers-lfs116/) - Cybercrime Investigations, Maltego: subject area published by the vendor, not a record of completion (https://www.maltego.com/categories/cybercrime-investigations/) ## Engagement shape Set out 2026-08-09. Four stages; most engagements do not run all of them. ### Framing Duration not fixed. One conversation before any scope is written, to establish what has already moved, what decision is waiting on it, and whether this is the right practice for the question at all. Part of that conversation is about the organisation rather than the estate: who would have to concede something for the likely recommendation to proceed, and whether that is currently a safe thing for them to do. A recommendation nobody can act on is a more common outcome than a wrong one. What the client receives: A written statement of the question as he understands it, which is often the first time it has been written down in one place. Where the honest answer is that someone else is better placed, that is said rather than worked around. ### Assessment Two to four weeks. Establishing what has actually changed, what is now exposed, and what can be measured. Risk engagements open here rather than with an implementation plan, because an implementation plan written before this stage is a guess with a timeline attached. What the client receives: A written assessment: the findings, the evidence for each one, and an explicit list of what could not be established in the time. That last list is the part most often left out, and it is the part that tells a board how much weight the rest will bear. Where an assessment returns a suspiciously short list of problems, that is reported as a finding about reporting rather than as a clean result. ### Architecture and decision support Duration not fixed. Turning the assessment into something a board or an engineering team can act on: the options, what each will cost beyond the licence, and which parts should be decided rather than delegated. What the client receives: Written architecture or a decision paper, in the register the audience actually reads. The doctrine documents on this site are published specimens of that writing, so the standard can be judged before it is commissioned. ### Handover Duration not fixed. Implementation and managed service are not undertaken here, so the work ends with material a client’s own team or supplier can execute against, rather than with a dependency on the person who wrote it. What the client receives: The documents, the reasoning behind them, and a named list of the assumptions that would have to be revisited if the situation changes. An engagement that cannot say what would falsify its own conclusions has not finished. Fees: Fees are settled once scope is agreed and are not published here, because the stages above differ too much in scope for any single rate to be honest across all of them. Not taken on: - Implementation or managed service delivery. Strategy, assessment and architecture only. - Expert witness or litigation support work. - Engagements where the conclusion has already been reached and an independent name is wanted for it. - Introductions sold as a service. Where a relationship is useful to a client it is offered, not invoiced. ## Notes archive ### The use of AI in combatting financial crime 2025-09-01. Panellist at a session convened by Standard Chartered’s Financial Crime Surveillance Operations, on the opportunities and pitfalls of applying artificial intelligence to financial crime detection. Alongside speakers from Deloitte and UEM Sunrise. ### On-demand learning on AirAsia Academy 2023-02-03. A short course on the fundamentals of entrepreneurship, leading change, and doing the right thing the right way. ### 47th ASEAN–Japan Business Meeting 2022-03-17. Presentation: accelerating social impact leveraging the metaverse, with a focus on the informal sector across ASEAN. ### Yayasan Peneraju Talk Series: jobs of the future 2020-12-09. A broadcast discussion on how tomorrow’s jobs will differ from today’s, and what learners should do to prepare. ### Digital transformation as corporate responsibility 2020-10-16. Delivered at a joint UNCTAD and UNITAR webinar on the role of entrepreneurship in post-pandemic recovery. Argues that governments should set standards and clear guidelines before structuring incentives. ### Are Malaysian companies ready for cloud adoption? 2020-10-02. Survey findings presented at the inaugural KL Cloud Maestro Series. An online survey of 105 respondents, average participating organisation about 1,200 people, led by Alphaus Cloud with Hitachi Sunway. Half of respondents cited missing skill sets as the largest gap after adoption; four in five named scalability as the principal benefit. The surviving recording is a short extract and shows the summary slide only. ### UNCTAD Multi-year Expert Meeting on Investment, Innovation and Entrepreneurship, eighth session 2020-09-22. Official summary statement to UNCTAD on accelerated digital adoption and the skills gap as the determining factor in recovery. ### Cloud computing for business leaders 2020-08-26. Hosted by the Malaysia–Japan Economic Association with the Malaysia South-South Association. On examining business models before they stop being viable. ### Community, foreign direct investment and the smart city 2020-07-24. A session hosted by Cyberview on driving a technology hub ecosystem through community rather than incentive alone. ### Tech This Way: an interview 2020-07-11. On first-hand experience across entrepreneurship, government and corporate roles, and the common challenges that recur in all three. ### ASEAN–Japan Business Meeting 2019-12-17. Conversations on startup ecosystem development, the automation of everything as a response to labour shortage, and mobility data. ### Yamato at one hundred 2019-12-14. On keeping a company relevant across a century, and the parcel delivery business that saved one from bankruptcy in 1976. ### Mentoring at Seedstars Summit Asia 2019-12-01. Facilitating workshop discussion on failures, lessons and what founders should understand about corporate partners before approaching them. ### Keynote at the RADIA inaugural event 2019-07-24. On access to the Japanese market for Malaysian businesses. ### A conversation on corporate turnaround 2019-06-23. Dinner in Tokyo with the chairman of Skymark Airlines, a founding partner of Integral Corporation and a well-known turnaround specialist. ### The digital divide in emerging economies 2019-04-02. Delivered at the United Nations in Geneva. On infrastructure, digital literacy, and the trade-offs between open markets and domestic digital growth. ### Western management, Asian wisdom 2018-04-13. Notes from the eFounders Fellowship in Hangzhou. On information as the new dividing line, giving a venture ten years, and the four years of failure behind a shopping festival. ### Keynote at the first Global Ventures Summit 2017-04-21. To global investors and startups, on connecting Southeast Asian technology companies with Silicon Valley capital. ### IFN Investor Forum 2016-04-15. On the Islamic funds landscape in the age of financial technology: blockchain, digital currencies and digital distribution.