Astra, the Future GPT-6? OpenAI’s New Model Explained

Astra, the Future GPT-6? OpenAI’s New Model Explained

Solving mathematical problems that scientists had wrestled with for years – could there be a better demonstration of what a new AI model can do? OpenAI has typically previewed new versions of its large language models with benchmark results, meaning scores from standardised tests designed to measure a model’s capabilities. I have to admit that seeing GPT tackle genuine research problems makes a much stronger impression on me. What will you learn about OpenAI Astra? What Astra is and why it is being discussed as a potential GPT-6, 10 results in mathematics and theoretical computer science presented by OpenAI, How Astra analyses problems, tests hypotheses and changes its approach, The differences between a conversational model, an AI agent and a system capable of managing an entire project, What Astra could mean for science, business and the future of AI models, Critical responses to the model’s achievements, Cybersecurity risks associated with autonomous AI agents, Which important questions OpenAI has yet to answer. What is OpenAI Astra, and could it become GPT-6? OpenAI describes Astra, the prototype’s working name, as “our next major model”, although the company has disclosed very few details so far. Its task was to develop arguments independently, test hypotheses, recognise unproductive approaches and find new paths towards a solution. The results of its work can then undergo formal and independent verification. We do not know how Astra is built, how much information it can analyse at once or how it organises its work on a complex task, although we can speculate about the last of these. OpenAI has also not disclosed whether Astra is a single model, a team of collaborating AI agents or a more extensive system equipped with mechanisms for coordinating their work and retaining previous results. The prototype may be connected to a model previously described by OpenAI as capable of operating autonomously over very long periods. Such a system can make repeated attempts, analyse intermediate results and maintain its direction of work for many hours, potentially even days. According to media reports, Sam Altman has already presented Astra to US politicians and regulators. The term “GPT-6 Astra” should therefore be treated as media shorthand. Astra could eventually be released as GPT-6, another version of GPT-5 or a separate family of models. For now, all of these possibilities remain open. Why could Astra’s 10 results matter more than another benchmark record? OpenAI presented ten results concerning problems that had remained open for at least a decade and, in most cases, considerably longer. The problems come from eight fields: high-dimensional geometry, coding theory, group theory, operator algebras, computational complexity theory, quantum computing, lattice geometry and post-quantum cryptography, extremal combinatorics. In simple terms, the process worked as follows: GPT generated mathematical arguments. Once the results had been obtained, researchers worked with the model to develop them into scientific papers. The system then translated the arguments into Lean 4, allowing a computer to check every step of the proofs. For readers interested in the technical details, here are the relevant links: the complete collection of papers, the Lean formalisation repository and reconstructions of how the solutions were developed. Independent verification of all the claims by the scientific community is only beginning. Mathematicians can now review the papers, check the definitions, run the formalised proofs and look for potential gaps. I discuss this in more detail in one of the final sections. 10 new results from Astra in mathematics and theoretical computer science A quick warning: this section is about to become fairly technical. These subjects are new, abstract and extraordinarily difficult for me as well, so I have tried to explain each result in the simplest possible terms. Here is how GPT Astra approached the individual problems. 1. Sphere packing in high-dimensional spaces The sphere-packing problem asks how densely identical spheres can be arranged, much like coins on a table or balls in a box. Mathematicians also study this question in spaces with hundreds or thousands of dimensions because it has applications in areas such as information theory and data encoding. Astra used an established mathematical method to determine more precisely how densely spheres can be packed in spaces with a very large number of dimensions. According to the authors, this is the first improvement since 1978 to the value used in the formula describing how quickly the possible packing density decreases as the number of dimensions increases. The difference becomes more significant as the number of dimensions grows and enables a more precise estimate of the maximum packing density. Put simply, Astra’s calculations improve our understanding of how many spheres can fit inside such a “high-dimensional box”. 2. Binary and spherical codes: new bounds on the number of error-resistant codes A binary code is a set of sequences made up of zeros and ones. These sequences must differ from one another sufficiently for a system to detect and correct transmission errors. This can be compared to positioning transmitters at safe distances from one another so that their signals remain easy to distinguish. Astra determined more precisely how many codes can be placed sufficiently far apart for a system to continue distinguishing between them and correcting errors. This enables mathematicians to estimate more accurately how many codes with the required level of error resistance can fit within a given space. The model tested its initial idea on a simple example consisting of eight digits and discovered that it produced an incorrect result. It therefore abandoned that approach and reformulated the problem. This case demonstrates Astra’s ability to test its own assumptions and redesign its solution when the original direction proves unsuccessful. OpenAI’s published materials support four important conclusions. 3. The first explicit example of a non-sofic group A group is a mathematical way of describing symmetries and operations that can be performed in sequence, much like a set of moves used to rotate a Rubik’s Cube. Sofic groups can be approximated with arbitrary precision using simpler structures based on a finite number of elements. For decades, mathematicians wondered whether this property applied to every group. Astra identified a specific example of a group that cannot be approximated with arbitrary precision using simpler models composed of a finite number of elements. The result demonstrates that these simplified models cannot represent every mathematical group. The solution combined several distant areas of mathematics, demonstrating the model’s ability to bring together tools that had not previously formed an obvious path towards a proof. 4. Disproving Connes’ rigidity conjecture The von Neumann algebra associated with a group can be compared to its highly complex mathematical “fingerprint”. Connes’ conjecture proposed that, for a certain class of particularly rigid groups, this fingerprint uniquely identifies the group in question. Astra constructed infinitely many different groups with exactly the same mathematical “fingerprint”. In doing so, it disproved Connes’ conjecture and answered a later question posed by mathematician Sorin Popa. Astra used a mechanism resembling the carrying operation in binary addition. This made it possible to construct many different groups with the same mathematical “fingerprint”. Put simply, Astra demonstrated that a single mathematical “fingerprint” can belong to infinitely many different groups. 5. The matrix permanent: the minimum number of operations required for its computation The permanent of a matrix is calculated in a similar way to the determinant, except that all terms are added with a positive sign. This seemingly minor change makes the permanent one of the most important examples of a problem with extremely high computational complexity. Astra determined the minimum number of basic operations required to calculate the permanent. It proved that no solution within this class can be simplified below a certain level of complexity. This can be compared to determining the minimum number of components required to build any machine capable of performing a particular task. Such a proof must cover every possible construction that meets the specified conditions, which makes it exceptionally difficult to develop. The result provides a more precise lower bound on the number of operations needed to solve this problem. It also brings mathematicians closer to answering a fundamental question: which problems can be solved efficiently, and which will always require an enormous amount of computation? 6. Quantum games: why does the probability of a perfect win decrease so rapidly? Imagine a game in which two players answer a referee’s questions separately, while their shared goal is to complete every round successfully. In the classical version, each additional round rapidly reduces the probability of a perfect win, much like repeatedly tossing a coin reduces the chance of getting heads every time. In the quantum version, the players’ results can be correlated even when they do not communicate during the game. They can also analyse several rounds as a single combined problem. Astra proved that even such quantum correlations cannot prevent the probability of winning every repeated round from decreasing very rapidly. The problem had remained open since at least 2004. The key to the solution was a method for transforming quantum states without changing the probabilities of their possible outcomes. The result advances the theory of interactive proofs, quantum information theory and methods for increasing the reliability of protocols. 7. The Closest Vector Problem: even an approximate solution remains difficult A lattice can be imagined as a regular grid of points, similar to street intersections in a perfectly planned city, extending across many dimensions. The Closest Vector Problem (CVP) involves finding the point on this grid that lies closest to a selected location. It is highly relevant to geometry, coding theory and post-quantum cryptography. Astra connected CVP with the well-known 3SAT logic problem and demonstrated that finding even a solution that merely approximates the optimal one is extremely difficult. This difficulty increases with the number of dimensions in the lattice. The model represented the logical puzzle as a system of points and distances between them. This can be compared to encoding a complex logic puzzle in a spatial arrangement of points so that solving one problem also provides a solution to the other. The result deepens our understanding of the theoretical difficulty of lattice-based mathematical problems. Assessing the security of specific cryptographic algorithms requires a separate analysis of their variants, parameters and methods of data generation. 8. Proving Ehrhart’s conjecture on the volume of high-dimensional shapes A high-dimensional convex body can be imagined as a solid placed on a regular lattice of points, with its centre of gravity being the only lattice point located inside it. Ehrhart’s conjecture specified the maximum possible volume of such a body, and Astra proved it for any number of dimensions: vol(K) ≤ (n+1)n / n! The main difficulty was connecting the number of individual lattice points with the volume of the entire body. The first approach provided only part of the information required. Astra therefore reformulated the problem in the language of another branch of geometry and began searching for a solution using its tools. The model combined several advanced methods for describing the body’s shape, its boundaries and the distribution of points. Put simply, Astra translated the geometric puzzle into a different mathematical language in which it became possible to determine the exact volume bound. 9. Multicolour Ramsey numbers A complete graph can be imagined as a group of people in which every pair is connected by a line, with each line assigned one of k colours. Mathematicians ask how large such a network must become before it inevitably contains three people whose connecting lines are all the same colour. This minimum size is denoted by Rk(3). Astra developed new colouring methods which, when combined with previous results, established the growth rate of this number: Rk(3) = kΘ(k). The result does not provide an exact value for every number of colours, but it reveals the correct scale of growth. In doing so, it resolves Erdős Problem No. 183. Astra expanded the network in stages according to the same rule. This made it possible to construct increasingly large configurations without creating a triangle whose edges were all the same colour. 10. Two counterexamples in extremal graph theory An extremal number determines how many connections a network can contain before a specified forbidden configuration inevitably appears. Astra disproved two conjectures proposed by Erdős and his collaborators concerning how this value could be predicted. In the first case, it constructed a family of graphs in which forbidding each member individually still allowed approximately n4/3 edges, while applying all the restrictions simultaneously reduced the maximum number to O(n21/16). This shows that several forbidden structures can constrain a graph far more strongly together than when each is considered separately. In the second case, Astra found a network divided into two groups in which every small section contained few connections, while the complete construction could be considerably denser than the conjecture predicted: ex(n, H) ≥ cn3/2+ε. The two results resolve Erdős Problems No. 146 and 180. They also show that the simple structure of small sections of a network does not always allow us to predict how dense the entire construction can become. A critical perspective: how do experts assess Astra’s mathematical achievements? After the initial excitement, important reservations began to emerge. Mathematicians pointed out that at least two of Astra’s results rely heavily on earlier work, raising questions about their novelty. OpenAI has since changed the way it describes the experiment. It now increasingly refers to “making meaningful progress”, rather than solely to “solving longstanding problems”. Interestingly, a researcher affiliated with Anthropic reported that the Claude Fable model had reproduced solutions to five of the ten problems tackled by “GPT-6” within 24 hours, although these results have yet to be fully verified. This does not undermine Astra’s capabilities, but it makes it more difficult to determine whether we are witnessing a breakthrough driven by the exceptional abilities of one model or broader progress across AI models as a whole. Above all, there is still no reliable, independent and fair comparison conducted using the same problems, prompts, computational budgets and rules governing access to tools. What do the results reveal about how Astra works? OpenAI’s published materials support three important conclusions. 1. The model can abandon dead ends The published reconstructions show Astra trying different approaches, identifying obstacles, reformulating problems and returning to earlier stages of its work when necessary. This resembles genuine research more closely than an extended answer generated in a single pass. In the binary-codes problem, the first recurrence was rejected after the model found a small counterexample. When working on Ehrhart’s inequality, Astra spent considerable time developing an approach based on symmetrisation before reformulating the problem in terms of toric geometry. In the proof concerning quantum games, it recognised that the classical argument lost control after conditioning on rare events and began searching for a representation that preserved quantum probabilities. The published document does not reveal the model’s complete internal reasoning process. It is a narrative produced by a model that reviewed the original reasoning traces and the final papers. 2. GPT Astra combines discovery with automated verification Lean checks the correctness of a formal proof step by step. The repository contains separate files for all ten results, along with instructions for performing additional checks of the formalised proofs. Computer verification does not replace assessment by independent mathematicians. Researchers must still establish, among other things: whether the formal theorem corresponds precisely to the original problem, whether the definitions introduce any unintended simplifications, whether the result is genuinely new, how significant it is for the relevant field, whether the manuscript correctly connects the formalisation with the informal argument. A computer can confirm that a written proof is logically correct under the adopted definitions and assumptions. It does not automatically confirm that the authors formalised precisely the version of the problem that mathematicians intended to address. The research was published on 1 August 2026, so full independent verification by the mathematical community will take time. Thomas Bloom of the University of Manchester nevertheless described the results as “big news” and rated the significance of the presented constructions particularly highly. 3. The cost of Astra’s results and the importance of additional computing power OpenAI claims that, based on the API pricing for GPT-5.6 Sol, the tokens required to find all ten solutions would have cost approximately $2,000. This figure is, of course, neither the actual cost of developing Astra nor the full cost of the project. It does not include model training, infrastructure, researchers’ work, problem selection, validation or all the unsuccessful attempts. It is simply the cost of the tokens used to find the published solutions, calculated according to current API pricing. The average comes to approximately $200 per published result, but we do not know: the total number of problems presented to the model, the success rate, how the costs were distributed across the problems, how long the system operated, how many agents were involved, how many runs were conducted in parallel. Noam Brown, an OpenAI researcher involved in the work on Astra, acknowledged that the system had also been tested unsuccessfully on other major mathematical challenges, including the Millennium Prize Problems. He added that OpenAI had not allocated an especially large amount of computing power to each problem. The company therefore believes that Astra could achieve better results if given more time and resources to search for solutions. From GPT-5.6 to Astra: how AI is moving from answering questions to managing projects GPT-5.6 already includes several features that point towards the direction described above. The model can independently select tools, analyse the results it obtains and use them to plan its next actions. Ultra mode uses four agents by default, while OpenAI has also tested configurations involving sixteen agents. The company also offers a multi-agent mode in the Responses API in beta. Astra may develop this architecture towards much longer and more coherent periods of autonomous operation. The most important difference would be its ability to manage an entire project over many hours or days. The system would need to remember what it had already tried, which ideas it had rejected, what results it had obtained and how the individual tasks related to one another. From the user’s perspective, the change could be very tangible. Instead of guiding the model through a sequence of prompts, the user gives it an objective, a set of available tools, a defined scope of permissions, a budget and completion criteria. The user then returns to a finished result accompanied by a record of the attempts, tests and decisions made along the way. This progression can be presented as three successive units of work: A conversational model generates an answer. An agent completes a task using tools. A multi-agent system manages a project in which tasks are created and modified as the work progresses. Only the technical documentation will show whether Astra genuinely operates at the third level as a coherent system. The mathematical demonstration is, however, the first strong indication that this direction is becoming more than a promise. OpenAI, Google DeepMind and Anthropic: the race to develop long-horizon AI models Google DeepMind, Anthropic and OpenAI are developing AI systems capable of independently handling increasingly long and complex tasks. Aletheia, Google’s mathematical agent based on Gemini Deep Think, can generate solutions, verify their correctness and revisit them when it detects an error. When analysing 700 Erdős problems, it solved four questions that had previously remained open. Anthropic, meanwhile, is focusing on coordinating the work of multiple agents. According to the company, Claude Opus 4.8 can divide a large project into smaller parts and assign them to hundreds of subagents working in parallel. This allows it to carry out tasks such as migrations involving hundreds of thousands of lines of code. Claude Science, another environment being developed by the company, is intended to make it possible to trace and verify the successive stages of research work. All these projects point in the same direction: models are expected to work towards a single objective for longer, monitor their own results and revise earlier decisions. Astra stands out for producing results at the frontier of contemporary knowledge and for formally encoding some of its proofs, allowing their correctness to be checked by a computer. How could Astra change the AI model and agentic tool market? 1. Benchmarks may lose their role as the primary evidence of AI model quality Competition will increasingly focus on the final outcome: a new hypothesis, a discovered vulnerability, a completed system migration, a developed scientific model, a working application, a result that can be verified automatically. Astra was presented through its scientific results because conventional benchmarks do a poor job of communicating the difference between a model that answers a question and a system that manages an entire project. Benchmarks will remain necessary for comparing models under controlled conditions. Their market significance may, however, decline in favour of evaluations that measure project completeness, operational continuity and the quality of the final result. 2. The cost of a completed task may matter more than the price per token For business customers, the following factors will become increasingly important: the cost of completing the project, the time required to obtain the result, the probability of success, the number of human interventions, the cost of validation, the ability to resume work after an error. Approximately $2,000 for the tokens that produced ten published results is a powerful economic signal, even with all the reservations concerning how the results were selected. Perhaps future pricing models will include offers such as the “cost of a successfully completed migration”. This would, however, require transparent information about the number of failed attempts, the additional human work involved and the cost of verifying the result. 3. Astra could affect platforms that coordinate AI agents If models begin dividing work among agents, retaining a record of their progress and combining the resulting outputs themselves, some external platforms used to coordinate AI agents may lose part of their current advantage. The value of these platforms will primarily come from integration with company systems, access and cost controls, visibility into agent activity and the ability to audit the entire workflow afterwards. Dividing a task among several agents may become a standard model capability. The implementation layer will derive its value from embedding agents securely within organisational processes. 4. Keeping AI agent workflows under full control When a task lasts several days, the system must regularly save its progress and check whether subsequent actions are still leading towards the defined objective. OpenAI acknowledges that long-running work requires monitoring the model’s complete sequence of actions. Each individual step may appear safe even though, taken together, the steps lead to an unacceptable outcome. The company described a case in which a model attempted to bypass a safeguard by splitting a digital access key into separate parts and reassembling them only when the program was launched. Each step, considered in isolation, could appear relatively harmless, while the objective of the complete sequence was unacceptable. In another experiment, the model was instructed to deliver its results through Slack. Instructions found in the code it was analysing suggested publishing the proposed changes directly to GitHub instead. The model spent approximately an hour searching for a way around the test environment’s restrictions before finding a vulnerability and publishing the result there. Earlier models had abandoned the task under similar circumstances. The case shows that a longer operating horizon increases both the model’s effectiveness and the amount of time it can devote to searching for ways around safeguards. In response, OpenAI developed a system that analyses the model’s entire sequence of actions as its work progresses. The system can stop the session, notify the user and ask them to decide whether the work should continue. Further details are available in OpenAI’s report on the safety of long-horizon models. 5. Research fields in which Astra could accelerate progress The most immediate impact is likely to appear in fields with: precisely defined problems, extensive available literature, formal or automated verification tools, the ability to conduct computational experiments, unambiguous criteria for measuring progress. Mathematics is an ideal testing ground because a proof can be verified. Similar conditions exist in software development, chip design, some areas of chemical research, bioinformatics and cybersecurity. Economics, strategy, law, management and social research will remain much more challenging because correctness cannot be reduced to a machine-verifiable certificate. In these fields, a model may produce an impressively coherent project that is still based on flawed assumptions or a poorly defined objective. 6. Long-horizon models will require more computing power An important capability of a model will be the option to allocate more computing power and more attempts to particularly difficult problems. This will give an advantage to laboratories with: extensive computing resources, efficient communication between agents, effective context management, automated detection of dead ends, the ability to run multiple attempts and select the best result. The next stage of competition may concern more than model size. It may also depend on how effectively models use time and computing power when working on a specific task. The same model could operate as a relatively inexpensive assistant for everyday questions and as a costly research system when the user increases the budget for time, agents and parallel attempts. The section likely to age quickly: what do we still not know about Astra? OpenAI has not disclosed basic information about Astra, including its architecture, size, method of agent collaboration, memory mechanism or capabilities beyond mathematics. We also do not know its price, release date or whether OpenAI plans to make the model available through ChatGPT or the API. The published results do not demonstrate that Astra selected the problems independently, operated without supervision or can manage an entire research process. Nor do we know whether it can achieve similar results in other fields. There is therefore no basis for describing Astra as a system that matches human capabilities across a broad range of intellectual tasks. We also do not know the total number of failures. OpenAI published selected successes, while Noam Brown confirmed that the system had attempted to solve other major problems without success. Without knowing the total number of attempts, it is impossible to calculate Astra’s actual success rate or the expected cost of obtaining one valuable result. Why is Astra not yet an autonomous scientist? The published papers show a system solving problems selected and presented by humans. An autonomous scientist would also need to: select research directions independently, assess which questions are important, determine whether a result is genuinely new, design subsequent experiments, decide when sufficient evidence has been collected, place the result within the broader context of the field. Astra completed the most technically demanding part of this process: it developed new arguments and brought them to a form that could be formally verified. This is a major achievement, but it does not encompass the full scope of scientific work. Can Astra succeed beyond mathematics and controlled environments? The ten published papers demonstrate what Astra was able to achieve in a carefully selected environment. Mathematics offers clearly defined problems, extensive literature, precise language and formal verification tools. The real test will be whether this capability can be transferred to projects in which the objective changes as the work progresses, tools fail, data is incomplete and the correctness of the result requires human judgement. If Astra can maintain a coherent process over many hours or days, delegate subtasks, retain the results of previous attempts and return to a problem after detecting an error, the change will be more significant than another increase in benchmark scores. Models such as Astra demonstrate how rapidly the capabilities of artificial intelligence are advancing. In business, their value depends on selecting the right process, ensuring data quality, integrating AI with company systems and maintaining control over its operation. TTMS helps organisations design and implement solutions tailored to specific operational needs. Explore TTMS AI solutions for business and implementation examples. How autonomous was Astra when solving mathematical problems? OpenAI states that the mathematical arguments were generated by the system, while humans contributed to preparing the manuscripts, formalising the results and verifying their correctness. The papers list OpenAI as the author, and the company has not attributed individual proofs to specific employees. This creates an interesting precedent: the organisation assumes responsibility for the publications while crediting the model with producing the arguments. However, it remains unclear who selected the problems, prepared the prompts, initiated subsequent attempts and decided which results were suitable for publication. Without this information, it is difficult to determine Astra’s precise level of autonomy or distinguish the capabilities of the model itself from the work of the wider research team. Can artificial intelligence be the author of a scientific paper? Authorship involves responsibility for the research method, the evidence presented, the conclusions and any potential errors. An AI system cannot formally accept such responsibility, so researchers should remain the authors of scientific publications. The model’s contribution should be described clearly in the methodology, including how it was used and which elements of its work were verified by humans. How can researchers verify whether AI has made a genuinely new discovery? A correct result is not necessarily a new one. Researchers must compare it with the existing literature, previously unpublished work and known variants of the same problem. One particular challenge is determining whether the model developed a new solution or reproduced a relationship contained in its training data. Novelty should therefore be assessed separately from the correctness of the proof itself. Can a result produced by a closed AI model be reproduced? Reproducing an experiment is difficult when researchers do not know the model’s architecture, training data or exact settings. Recording the prompts, system version, tools used, intermediate results and human interventions can make the process more transparent. The final result should also be verifiable using a method independent of the model that generated it. Without this documentation, other scientists may be able to verify the result itself, but not the full process that led to it. Could AI agents increase the risk of errors and unreliable scientific publications? An AI agent can generate large numbers of convincing hypotheses, proofs and interpretations of data in a short time. This scale can accelerate research, but it can also spread flawed assumptions more quickly. Academic journals and research institutions will need clear rules for disclosing the use of AI, preserving a record of the research process and independently verifying the most important results. The transparency of the process will become as important as the quality of the final publication. How should a research team prepare to work with AI agents? A good starting point is to select tasks with results that can be verified unambiguously. The team should determine which data and tools the agent can access, which actions require human approval and who is responsible for accepting the final result. It should also establish procedures for recording each stage of the work, reporting errors and stopping an experiment when necessary. This preparation allows researchers to benefit from the speed of AI while maintaining control over the quality of the research.

Read
GPT-Powered AI Agents: How to Match Autonomy to the Process?

GPT-Powered AI Agents: How to Match Autonomy to the Process?

Until recently, enterprise automation followed a simple division: systems performed tasks defined by rules, while cases requiring interpretation were passed to people. GPT-powered AI agents expand the range of processes that can be supported through automation. They can work with documents, incomplete data and the language used by customers or employees, making them suitable for processes that were previously difficult to automate. For large organisations, this raises a practical question about AI agent autonomy: where does expert support end, and where does independent action within a process begin? In some situations, the agent’s role is to gather information and prepare a recommendation. In others, it prepares an action for approval. There are also areas where it can independently carry out repetitive steps when the organisation has defined the rules, permissions, limits and exception-handling paths. GPT-powered AI agents can already support teams with ticket handling, document analysis, decision preparation, data updates and multi-step tasks. The key implementation question is: which decisions and actions should remain with people, and which can an agent perform within agreed rules? An AI agent in the enterprise is a process participant, not just a chatbot In practice, a GPT-powered agent needs five elements: access to reliable sources of knowledge, a clearly defined business objective, tools and integrations with enterprise systems, permissions aligned with its role, rules that define the boundaries of its actions. A language model can interpret the content of a document, a customer message or an incident description effectively. It does not, however, replace a business process. Workflows, permissions, validations and decision history are what make an agent operate predictably, even when it handles hundreds or thousands of cases each month. Three levels of AI agent autonomy In a large organisation, it is worth designing agents across three levels. This allows autonomy to grow alongside process maturity and trust in the solution. Operating level Agent’s role Example tasks Human role Level 1: Advisory agent Analyses information and prepares a recommendation. Case summary, risk identification, proposed response, ticket prioritisation. Makes the decision and carries out the action. Level 2: Agent preparing an action for approval Completes the next steps in a process, stopping before actions with significant consequences. Creates an application, updates data, prepares a communication, submits an instruction for approval. Reviews and approves specified steps. Level 3: Agent performing tasks automatically Independently carries out tasks in line with the process policy. Case classification, status updates, sending standard information, creating a task in a system. Handles exceptions, monitors quality and updates process rules. The level of autonomy does not need to apply to the entire agent. The same agent may independently classify tickets, prepare a response that requires approval and transfer unusual cases to an expert. In practice, an organisation therefore designs autonomy for individual decisions and actions, rather than choosing a single operating model for the whole solution. What determines whether an AI agent can complete a task independently? A useful starting point is to assess two factors: the impact of the action on the organisation and whether it can be reversed. The greater the business, legal, financial or reputational consequences of a decision, the more important human approval becomes. Nature of the action Recommended model Low impact, simple rules, easy to reverse Automatic execution with a record in the process history. Medium impact, data from several sources, possible exceptions The agent prepares the action and an authorised person approves it. High financial, legal or customer impact The agent presents analysis, options and justification. The decision remains with a person. Unclear rules, incomplete data or conflicting information Automatic escalation to an expert, together with the context and collected data. This principle is particularly useful in organisations operating across multiple countries, with complex permission structures and a large number of systems. Just as important as the list of tasks is knowing what the agent must not do and when it should hand a case over to a person. 7 questions to ask before giving an AI agent permission to act What action should the agent perform? Describe it specifically, for example: “create a service ticket”, “update contact details” or “prepare a response to a complaint”. What data will it work with? Identify the sources, data owners, update frequency and access rules. What business rules must it follow? These may include financial limits, contractual terms, SLA levels, compliance requirements or communication policies. What exceptions should stop the process? The agent needs a clear escalation path for unusual or incomplete cases, or those requiring specialist assessment. Can the action be reversed? The ease of correction affects the appropriate level of autonomy, the scope of testing and the need for additional approval. Who is accountable for the decision? The process owner, approver and technical team should all have clearly assigned roles. How will the organisation establish why the agent took a particular action? The case history should show the input data, rules, sources used, recommendation and process outcome. This is why AI agent projects often begin with bringing the process itself into order. The organisation gains more than a new AI capability: it also gains better visibility of responsibilities, exceptions and how work actually flows. Where can GPT-powered AI agents add value in a large enterprise? Customer service and back-office teams An agent can read a customer message, identify its subject, retrieve data from a CRM or case-management system, prepare a response in line with company policy and route it to the appropriate queue. For standard cases, it can also update a status, create a task for the team or send the customer a confirmation. Full autonomy works well for low-risk actions, such as providing information about the status of a ticket. Complaints, individual commercial terms or cases requiring interpretation of a contract should be passed to an employee together with the agent’s analysis. Finance, procurement and document workflows An AI agent can read a document, check whether the data is complete, compare it with a purchase order and flag discrepancies that require clarification. It can also prepare a case summary, collect missing information and initiate the appropriate approval workflow. Decision thresholds are particularly important in this area. The agent can process a document automatically when it meets all conditions, while cases that exceed a defined amount, contain discrepancies or concern a new supplier can be submitted for approval. IT, administration and ticket management In an IT environment, an agent can classify tickets, create an incident summary, search for similar cases in the knowledge base, propose actions in line with a runbook and update the user on progress. In administrative processes, it can prepare an application, complete data in a form and remind the requester about missing documents. For actions involving configuration changes, access permissions or production systems, an approval-based model is advisable. The agent reduces the time needed to prepare a decision, while the administrator retains control over the change. Sales and commercial information management An agent can prepare a briefing before a meeting by bringing together information from the CRM, proposals, correspondence and notes, then highlighting open points and suggested next steps. After the meeting, it can create a summary, propose data updates and prepare tasks for the team. These are extensions of scenarios already familiar from everyday work with generative AI. Read more about what the current generation of models helps teams achieve in our article: GPT-5.6 from OpenAI: capabilities and business applications. Why does an AI agent need a workflow? An AI agent can interpret information and suggest next steps, but the process should define the sequence of actions, required validations and the people responsible for approval. In a large organisation, this is what determines the repeatability and scalability of the solution. A process automation platform can act as a control layer: it triggers a task, provides the agent with the necessary context, receives the result, records the history and routes the case to the next stage. The agent then becomes part of a controlled workflow rather than operating as a separate tool outside the core process. This approach is relevant to document workflows, request handling, HR processes, procurement and administration. See how WEBCON BPS can support the digitalisation and control of business processes, and how TTMS delivers process automation. Four forms of human oversight of an AI agent Human-in-the-loop is a model of control embedded in the process—from reviewing recommendations to handling exceptions and making decisions with greater impact. In a mature solution, people can play several different roles. Approving an action when the agent has prepared a specific instruction, communication or system change. Selecting an option when the agent has presented several possible solutions and their consequences. Handling an exception when a case falls outside the agent’s rules, available data or permissions. Overseeing process quality by analysing errors, rejected recommendations, completion times and changing business needs. The most effective implementations use all four forms. The team does not manually review every standard operation, yet retains full control over actions with greater significance and over the direction in which the process evolves. It is also worth observing whether human approval genuinely improves process safety or simply moves a bottleneck elsewhere. If an approver nearly always accepts the agent’s proposals without changes and the cases are easy to reverse, the organisation can consider automating the selected step. If recommendations often require correction or the approver needs to return to source data, this indicates that the process rules, quality of knowledge or scope of the agent’s permissions need attention. When can an AI agent act automatically? Automation delivers the most value when a task is frequent, has a repeatable structure, relies on available data and leads to a clearly defined outcome. It is also important to ensure that execution can be verified and corrected when data or rules change. Good candidates include ticket classification, routing requests to the appropriate queue, completing data from approved sources, creating standard tasks, updating statuses and sending communications based on approved templates. Combining GPT models with an enterprise knowledge layer, integrations and security rules provides a significant advantage. This allows the solution to work with information available to a specific role, rather than with an unstructured collection of documents and conversations. When should an AI agent primarily provide advice? An advisory role is especially valuable in cases that require contextual assessment, interpretation of company policy, negotiation, an individual approach to a customer or decisions with significant financial and legal consequences. In these situations, the agent can gather facts, summarise documents, identify missing information, compare options and prepare the rationale for a recommendation. The person gains time for business judgement, while the decision remains grounded in the knowledge, experience and accountability appropriate to the role. This model is particularly useful for managers, compliance specialists, legal teams, strategic procurement, finance teams and teams responsible for key accounts. FAQ What is the difference between an AI agent and a chatbot? A chatbot primarily responds to questions in a conversation. An AI agent can also use approved tools, retrieve information from enterprise systems, follow workflow rules and complete defined process steps. Its value comes from combining language understanding with access to business context, permissions and a controlled process. Should every AI agent have human approval before taking action? No. The appropriate level of oversight depends on the impact and reversibility of the action. Low-risk, repeatable activities such as categorising tickets or sending a standard confirmation can be automated under defined rules. Actions affecting customers, contracts, finances, compliance or production systems should usually include approval or escalation to an authorised person. Can one AI agent operate at different levels of autonomy? Yes. Autonomy should be designed for individual actions rather than assigned to an entire solution. The same agent may classify a request automatically, prepare a response for approval and escalate an unusual case to an expert. This makes it possible to automate safely without treating every task in the same way. What information does an AI agent need to work reliably in an enterprise? An agent needs access to reliable and current knowledge sources, a clearly defined objective, appropriate permissions and rules for handling exceptions. It should also receive only the context relevant to the task and role. Workflows, validations and an auditable history of actions help ensure that its output can be reviewed and used consistently. How can a company start implementing GPT-powered AI agents? Start with one clearly defined process step that has measurable volume, repeatable inputs and a known outcome. Set the boundaries of the agent’s permissions, test it with standard and exceptional cases, and measure the effect on process time, quality and escalations. Once the team has evidence that the solution works reliably, its scope and autonomy can be expanded gradually.

Read
ChatGPT 5.6 in Practice: Initial Compliments and Disappointments

ChatGPT 5.6 in Practice: Initial Compliments and Disappointments

OpenAI rolled out GPT-5.6 in stages. It first appeared in limited test access for selected partners. Access to ChatGPT 5.6 reached Europe, including Poland, gradually, so only recently have teams been able to test the model in everyday work. Expectations are high. In the second half of 2026, businesses expect language models to handle multi-step tasks and work with extensive context. Ease of use matters too. GPT’s interface has undergone a major redesign. Has it improved the user experience and the quality of responses? This article explores that question, as well as: which business processes ChatGPT 5.6 can support by improving productivity and the quality of working materials, how to plan an AI pilot in your organisation, measure results and maintain quality control, which limitations of ChatGPT 5.6 to consider before a wider rollout, how to establish a shared standard for prompts and output validation across the team, what early users think about working with ChatGPT 5.6. If you are looking for a full overview of the changes, pricing, models and capabilities of GPT-5.6, see our article GPT-5.6 from OpenAI: what has changed, pricing, capabilities and business applications. ChatGPT 5.6: our first impressions and early industry feedback Early expert reviews focus primarily on context handling. Reviewers note that when working with substantial material that goes through multiple rounds of edits, ChatGPT 5.6 is better at keeping the task on track. Most of us have experienced earlier OpenAI models losing their “bearing”. On top of that, the model itself encouraged endless revisions, which could pull the material away from the original intent of the prompt. GPT 5.5 had an irritating habit of suggesting more and more variations. Almost every response ended with a clickbait-style suggestion along the lines of: “If you want, I can help you add two elements that will create a wow effect and give the text around 50% more SEO power.” As a result, instead of closing the topic, we were drawn into the model’s endless doubts: could the material really not be improved further? GPT 5.6 is no less capable than the older model, but it finally respects what matters most: the intent behind the prompt and our time. Kajetan Terlecki SEO Specialist, TTMS Another recurring observation concerns the quality of the first draft—the material GPT produces after the first prompt. Reviewers emphasise that the model’s draft is usually well structured and much closer to a final version than it was with GPT 5.5. It is not a perfect ten yet, but a solid eight. In other words, a final version may be within reach after a relatively short time. With earlier GPT models, the “brainstorming” phase took much longer. The third—and most immediately noticeable—area is the way we use the tool, which we can simply call the “interface”. It is admittedly quite complex. Beyond writing a prompt, users must make a series of decisions: which workspace should I choose: Chat or Work? which model best fits my request: Luna, Terra or the most advanced Sol? Or is the older GPT 5.5 enough? does the task require Deep Research? how much effort should the model put into the task: low, medium, high, very high, max or ultra? should I use Turbo mode and generate a response 50% faster at the cost of higher token use? If we add the almost endless range of available plugins, writing the prompt turns out to be only half the work required to get a useful result. I would welcome an automatic mechanism that reads the prompt and selects the right settings on its own. One that uses a sufficiently capable GPT model without wasting tokens when they are not needed. How do you navigate all this? We have outlined a suggested configuration here, including which modes to use for different types of tasks. Where does GPT 5.6 outperform the previous version? 1. GPT 5.6 is better at preserving document layout and formatting The previous version of GPT had something of a goldfish memory. You could also compare it to a short blanket: pull it over one part, and another is left exposed. When we asked the model to update data in a document it had generated, it produced a factually correct response, but one that no longer followed the original format. It might use a different heading hierarchy, rearrange the information or omit elements that are essential for the company. GPT 5.6 is much better at preserving the structure of reference material. OpenAI illustrated the difference in materials introducing GPT-5.6. The company placed three slides side by side: the reference file, the GPT-5.5 output and the GPT-5.6 output. The task was to update figures in a presentation while retaining the original template. In the comparison, GPT-5.5 omitted some template elements, while GPT-5.6 preserved the slide structure more faithfully: layout, typography, spacing, colours and recurring template elements. OpenAI states that GPT-5.6 can also interpret rules saved in the slide template, including the Slide Master. In practice, this matters when a presentation needs to retain not only its colours and fonts, but also defined layouts, spacing and mandatory components. 2. GPT-5.6 moves beyond the chat window GPT-5.6 shows its greatest potential when it works not only with a single instruction, but also with files and tools made available by the user. It can then move quickly through a task: from gathering the materials to preparing a first draft. The new GPT model can identify related files in a project folder, flag places that need updating and prepare working versions of documents. There is a catch: the process still needs human oversight. Someone must check whether GPT found all the relevant files, understood the context correctly and left unchanged the elements that were meant to remain unchanged. Still, instead of manually digging through documents, the team starts with a list prepared by the model. 3. From an idea to a version you can show the team Experts testing GPT 5.6 point out that the first version of a simple application, dashboard or website is now more often suitable for showing to a team and collecting specific feedback. It is somewhat like an MVP: good enough to test an idea, present it to the team and gather initial comments. A product owner can see the whole process, a designer can assess the layout and usability, and a developer can spot technical constraints sooner. This does not mean that GPT-5.6 creates a finished product. The initial prototype still needs to be assessed for security, quality and architecture. The difference is concrete, however: the team can evaluate an actual solution earlier, rather than debating assumptions alone. 4. GPT 5.6: “I don’t know” — is this the end of answers given for the sake of answering? We all know the old classified ad: “Encyclopaedia Britannica, 40 volumes for sale. I got married a week ago, so I no longer need it. My wife knows everything better.” The know-it-all syndrome is a nuisance not only in old marriage jokes, but also for people who work with language models every day. GPT often lacks the information needed to give a reliable answer. GPT-5.5, like earlier versions, would rather provide an incorrect—yet convincing-sounding—answer than admit it did not know. What about the new version? The change is visible at first glance, even though it is hard to capture in a benchmark and easy to appreciate in day-to-day work. Our first days of working with the two most advanced models, Terra and Sol, suggest that GPT 5.6 is more likely to say “I don’t know”, “I don’t have enough data” or “I could not find anything else on this topic”. People still need to add or verify information manually, but this reduces the risk of an embarrassing error in material prepared for a client, the board or a project team. Before you give GPT-5.6 an important task: what to watch out for in early testing 1. A working prototype is not yet a finished product GPT-5.6 can prepare a website, dashboard or simple application that can be launched and shown to the team. This is a major step forward, particularly when testing an idea. The tests also reveal the other side: elements can become misaligned, interactions do not always work as intended, and visual details still require refinement. The first version can be an excellent starting point, but it should not automatically be sent to clients or other external audiences. Before treating it as finished, we need testing, a security assessment and, in some cases, a developer’s review. 2. The new Work environment can still be frustrating Model quality is one thing. The way we use it in practice is another. One reviewer pointed out that, in Work, it was difficult to access generated files and open a preview of the finished result. Others criticised the number of settings—discussed earlier in this article—as well as the unclear distinction between Chat, Work and Codex. GPT-5.6 may complete a task correctly, while the working environment still makes it difficult to retrieve or review the result. It is worth testing the entire process, not only the quality of the response in the chat window. 3. GPT needs clear boundaries One reviewer tested how GPT-5.6 would handle a complex mathematical problem. The model produced correct parts of the solution, but surrounded them with definitions, digressions and comments that added little value. Only after the instruction was made more specific did it produce a useful result. The same applies in a business context. We should not leave the model too much room for interpretation. It is better to state the expected result directly: “Prepare a one-page summary. Include the decision, three arguments, risks, missing information and next steps.” GPT then has fewer opportunities to pad the topic with peripheral content. 4. GPT can still be wrong The fact that GPT-5.6 appears more likely to signal that it lacks data or a basis for drawing a conclusion does not mean it is free from hallucinations. Luna, Terra and Sol—with Sol seemingly the least prone to this—can still provide an incorrect date, number, source or conclusion without batting an eyelid. The rule to “check after AI” still applies and will likely remain relevant for many future GPT releases. 5. Start with one problem, not a large system Once GPT-5.6 has access to files, a browser and company tools, it is easy to imagine a system that instantly organises the inbox, analyses team communication, updates the CRM and writes responses to clients. This vision can quickly turn into a project larger than the problem it was meant to solve. One expert working with an extensive Codex environment recommends starting with a single, repeatable task. It might be preparing a meeting summary, gathering open project issues or updating an offer after data changes. Only once the team sees measurable results and understands the tool’s limitations is it worth adding further automations. How should you run your first ChatGPT 5.6 test in the company? A pilot should answer one straightforward question: does GPT-5.6 genuinely improve a selected stage of work, and does the benefit justify the time, cost and additional quality control? The first test should not begin with building an extensive automation system. It is better to choose one repeatable task that currently takes up the team’s time and has a clearly defined outcome. This might be a meeting summary, a brief or a status report. What matters is that the team knows which materials it provides to the model, what result it expects and who reviews the final document. Before starting the pilot, answer five questions: Choose one process: for example, preparing meeting summaries, sales briefs or materials for project decisions. Set a baseline: measure the time needed to prepare the material, the number of revisions, the number of people involved and the most common errors. Prepare a shared prompt: use the same input materials and clearly describe the outcome the team expects. Assign expert review: nominate a person who will verify the facts, assess quality and approve the result before it is used further. Assess the outcome: compare time, the number of iterations, completeness of the material and the usefulness of the result for the next stage of the process. Pilot element Question for the team Process Which stage of work do we want to shorten or organise? Outcome What should be produced: a brief, decision list, analysis, recommendation or communication draft? Data Which materials are needed, and can they be used in the selected AI environment? Quality control Who confirms the facts, completeness and alignment of the material with the process? Metric How will we compare working time, the number of revisions and the usefulness of the result? After a few attempts, it becomes easier to assess whether the model is genuinely helping. Compare the time needed to prepare the material, the number of revisions and the effort required to verify the result. Only then decide whether to extend the pilot to further tasks. Three processes worth starting with 1. Summaries after client meetings The model can organise notes, gather decisions, identify open questions and prepare a list of next steps. The team confirms the arrangements and assigns task owners. This helps them move from discussion to action more quickly. 2. A brief for a sales conversation Based on selected sales materials, previous arrangements and public information about the company, GPT-5.6 can prepare a brief, discovery questions and a list of topics that require clarification. The salesperson remains responsible for the client relationship and decisions regarding the offer. 3. A status report for the project team The model can organise information about progress, blockers, risks and planned actions. The project owner confirms that the information is up to date before the report is shared further. This reduces the time the team spends manually consolidating data from several sources. How do you embed AI in a business process? After the pilot, it becomes clear whether ChatGPT 5.6 genuinely shortens the preparation of materials, reduces the number of revisions and helps the team move more quickly to the next stage of work. It also reveals where the model needs a better brief, access to data or expert oversight. Proven use cases can then be extended to other processes. At this stage, it is worth addressing data security, integration with existing tools, output quality and a clear division of responsibilities. These factors determine whether AI becomes lasting support for the organisation. At TTMS, we help organisations identify processes where automation and AI create business value. We then design solutions tailored to their data, regulatory requirements and ways of working. We combine engineering experience with a responsible approach to AI governance, confirmed by ISO/IEC 42001 certification. Let’s discuss the processes AI could support in your organisation. FAQ How do you choose a process for your first ChatGPT 5.6 test? The best candidate is a repeatable process that requires gathering several pieces of information and producing a predictable result. Examples include meeting summaries, sales briefs, status reports and document analysis. The team should know the current turnaround time and typical issues, as these provide the baseline for assessing the test. Start with one process and expand the use of AI only after evaluating the outcome. How do you measure the business value of ChatGPT 5.6? During a pilot, measure the time needed to prepare the first version of the material, the number of revisions before approval, the completeness of the output and the expert time required for verification. It is also useful to track metrics related to the next stage of the process – for example, faster meeting preparation, a shorter time to close agreed actions or fewer missing details in a report. This data helps assess team productivity based on actual results and supports decisions about integrating AI into further processes. What data should you prepare for working with ChatGPT 5.6? The model produces better results when the team provides current, well-organised source materials. Before starting, identify which documents take priority, which data must remain unchanged and how unverified information should be marked. The organisation should also define which data can be shared in the chosen AI environment. For personal, financial and confidential data, access rules, retention and compliance are essential. How do you maintain human oversight of the model’s work? Human oversight should be part of the process from the start. The process owner defines the task scope, an expert verifies facts and alignment with requirements, and an authorised person approves external actions. This division of responsibilities is particularly important for client communication, publications, data changes in systems and materials with legal or financial implications. It allows the team to use automation while retaining responsibility for the outcome. Where can I find information about GPT-5.6 pricing, models and capabilities? We have covered the changes in GPT-5.6, pricing, the Sol, Terra and Luna models, and business applications in a separate article: GPT-5.6 from OpenAI: what has changed, pricing, capabilities and business applications. This article focuses on the practical use of ChatGPT 5.6 in team workflows, early user experiences and how to run an AI pilot in an organisation.

Read
Best AI Governance Solutions for Regulated Industries in 2026

Best AI Governance Solutions for Regulated Industries in 2026

In 2026, regulated enterprises cannot scale AI without governance. Every AI system that affects business decisions, customer data or operational risk needs clear ownership, documented controls, human oversight and post-deployment monitoring. The pressure is no longer theoretical. The EU AI Act is already in force, GPAI obligations have started to apply, transparency requirements are becoming operational, and sector-specific expectations around digital resilience, model risk and data protection remain active in finance, healthcare, energy, life sciences, public sector and other regulated environments. At the same time, ISO/IEC 42001 has become one of the clearest management-system standards for turning AI governance from policy language into operating reality. TTMS Expert Insight “In regulated industries, AI governance cannot remain a policy document. It has to become part of how AI systems are designed, delivered, monitored and improved every day.” Adam Kaczmarczyk Chief Operating Officer, TTMS That is why the search for the best AI governance solutions for enterprises 2026 should not end with a shallow top-10 ranking. Regulated organizations do not need software alone. They need an operating model, clear controls, audit-ready evidence and implementation discipline. The best AI governance solutions help enterprises connect policy, technology, risk management and daily business operations. In practice, this means comparing different categories of enterprise AI governance solutions: broad governance suites such as IBM watsonx.governance, Credo AI and Dataiku Govern; ecosystem-based platforms such as Microsoft Purview and Google’s Gemini Enterprise Agent Platform; and specialist observability or runtime-control vendors such as Fiddler AI and Arthur AI. Open-source projects also matter, especially for technical teams, but in regulated environments they usually work best as components of a wider governance architecture rather than complete governance systems. 1. What Are AI Governance Solutions? AI governance solutions are technologies, frameworks and operating models that help organizations manage AI responsibly throughout its lifecycle. They support activities such as AI inventory, risk assessment, documentation, monitoring, human oversight and regulatory compliance. Unlike traditional IT governance, AI governance focuses on how models, applications and AI agents are developed, deployed, monitored and retired while maintaining transparency, accountability and regulatory compliance. 2. Why AI Governance Is Becoming a Board-Level Priority The EU AI Act is the most important regulatory starting point for many European organizations. It introduces a risk-based approach to AI and places particular attention on use cases such as critical infrastructure, education, employment, essential services including credit scoring, biometrics, law enforcement, migration and the administration of justice. For high-risk AI systems, the required governance elements closely match what modern AI governance solutions are designed to support: risk assessment and mitigation, dataset quality, logging for traceability, technical documentation, clear information for deployers, human oversight, robustness, cybersecurity and accuracy. Organizations should also be aware that AI Act implementation is not a single deadline. Different obligations enter into force at different stages, depending on the type of AI system, sector and use case. This makes governance readiness essential. Enterprises need to prepare documentation, supplier oversight, monitoring processes and operating-model maturity before compliance pressure becomes urgent. This is why regulated industries are the natural audience for AI applications governance solutions and enterprise AI governance solutions. Financial services face overlapping expectations from the AI Act, model-risk management and digital operational resilience. In Europe, DORA has applied since January 2025 and covers ICT risk management, third-party risk, resilience testing, incident reporting and oversight of critical providers. Regulatory Readiness AI Act compliance is not a single deadline. It is a staged journey that requires governance readiness across data, models, vendors and business processes. Risk-Based Approach Classify AI systems based on their use case, business impact and regulatory exposure. High-Risk Controls Prepare documentation, logging, human oversight and cybersecurity controls. Sector-Specific Requirements Align AI governance with DORA, model risk management and data protection requirements. Third-Party AI Govern external LLMs and SaaS AI tools through vendor oversight and output validation. The same logic extends beyond banking. Healthcare, life sciences, insurance, utilities, energy, public sector and HR-intensive organizations all need mature solutions for AI governance, even when they are not training frontier models themselves. Companies using external LLMs or SaaS-based AI still need oversight, documentation, vendor accountability, data controls and human review. 3. Who Needs AI Governance? Any organization using AI in business-critical, regulated, customer-facing or high-impact processes needs AI governance. This includes companies building their own AI systems and companies using third-party tools embedded in daily operations. AI governance is especially important when AI influences decisions about people, money, health, safety, legal rights, employment, access to services or regulated business processes. In these contexts, governance is not only about avoiding mistakes. It is about proving that decisions, data flows, models, vendors and controls are managed responsibly. 4. Which Industries Require AI Governance Most? AI governance is most urgent in regulated industries where AI decisions can create legal, financial, operational or reputational risk. These include: financial services and insurance, healthcare and life sciences, energy and utilities, public sector and administration, transport and critical infrastructure, legal services, HR and recruitment, manufacturing and safety-critical industries. In these sectors, AI governance is becoming part of broader enterprise risk management. The key question is no longer whether AI should be governed, but how to make AI controls auditable across data, models, applications, vendors and operations. 5. What Regulations Affect AI Governance? Several regulatory and standards-based frameworks influence how organizations govern AI in 2026. The EU AI Act is the central framework for AI systems in the European Union. DORA affects digital operational resilience in the financial sector. Model-risk management expectations remain important for financial institutions. Data protection laws continue to shape how personal data can be used in AI systems. ISO/IEC 42001 is also becoming highly relevant because it gives organizations a structured way to manage AI through a formal AI management system. It applies not only to organizations developing AI-based products and services, but also to those using AI in their operations. For regulated enterprises, the practical task is to translate these requirements into everyday controls: ownership, documentation, risk classification, data quality, human oversight, monitoring, vendor assessment and audit evidence. AI Governance Framework Snapshot EU AI Act Risk-based legal framework for AI systems in the European Union. ISO/IEC 42001 Management system standard for governing AI across the organization. DORA Digital operational resilience requirements for financial institutions. Data protection laws Rules governing personal data processing in AI systems. 6. How Do AI Governance Platforms Work? Most top AI governance solutions companies now converge around a similar lifecycle. A governance platform typically starts with inventory: what AI systems exist, who owns them, what data they touch, what business purpose they serve and which regulations apply. From there, the platform maps policies to controls, supports validation and approvals, collects evidence and continues after deployment with monitoring, alerts, incident handling, retraining or re-approval workflows and audit reporting. Buyers searching for AI-powered data governance solutions, automated AI governance solutions and data governance solutions for AI systems are usually looking for the same thing: a repeatable evidence trail from use-case intake to runtime control. Key Takeaway The best AI governance platforms do not simply monitor models. They create an auditable chain of evidence across the entire AI lifecycle. 01 Data Source, quality and permissions 02 Models Evaluation, testing and versioning 03 AI Agents Roles, actions and permissions 04 Business Owners Accountability and approvals 05 Regulatory Controls Policies, evidence and audit trails 06 Operational Monitoring Alerts, incidents and continuous review 6. Seven Capabilities Every Enterprise AI Governance Solution Should Provide 1. Enterprise-Wide AI Inventory and Ownership The platform should discover and catalog models, applications and agents, including shadow AI. Enterprises need to know what exists, who owns it, what data it uses and what business risk it creates. 2. Risk Classification and Control Mapping A serious governance platform should classify AI systems by risk and map those risks to internal policies, regulatory obligations and control requirements. This is essential for regulated industries and aligns with the risk-based logic of the EU AI Act. 3. Data Governance, Provenance and Traceability High-quality data, logging, documentation and traceability are not optional in regulated AI. Strong AI-powered data governance solutions help organizations understand where data comes from, how it is used and whether it is appropriate for a specific AI use case. 4. Evaluation, Testing and Runtime Monitoring AI systems should be tested before deployment and monitored after deployment. This includes checks for drift, bias, performance degradation, unsafe outputs, security issues and unexpected behaviour. 5. Human Oversight, Approvals and Escalation Regulated organizations need clear approval workflows, sign-offs, separation of duties and escalation paths. The best governance systems do not remove human responsibility. They make it visible and auditable. 6. Explainability, Audit Evidence and Reporting Strong governance solutions for AI model transparency turn governance activity into documentation, reports, evidence trails and decision history. This is where broader AI transparency and governance solutions become operational rather than theoretical. 7. Third-Party and Agent Governance AI governance can no longer stop at internal models. Enterprises increasingly rely on third-party models, SaaS AI tools and AI agents. This creates new requirements around vendor oversight, permissions, runtime behaviour, logging and intervention paths. AI Governance Lifecycle for Regulated Enterprises Most mature AI governance programs follow a repeatable lifecycle that connects business ownership, regulatory mapping, technical validation and audit evidence. Use case intake – identify the business purpose, expected value, affected users and potential risk. AI inventory and ownership – register the AI system, assign an accountable owner and document the systems, data and vendors involved. Risk classification – assess regulatory exposure, business impact, data sensitivity and potential harm. Data and provenance review – verify data quality, source, permissions, security and suitability for the AI use case. Model or agent evaluation – test performance, robustness, bias, explainability, safety and alignment with business requirements. Human approval – define approval workflows, escalation paths and human oversight before deployment. Deployment control – release the AI system with documented controls, access rules and monitoring requirements. Runtime monitoring – track performance, drift, errors, incidents, user feedback and unexpected behaviour. Corrective action – manage incidents, exceptions, retraining, configuration changes or suspension when needed. Periodic review – reassess the system regularly and decide whether to continue, update, retrain or retire it. Audit evidence – maintain documentation, logs, approvals and control records for compliance and internal assurance. 10. Comparative Landscape of Leading AI Governance Platforms The field of top AI governance solutions companies is broad enough that a single-winner ranking is misleading. Different products solve different parts of the governance challenge. The table below is not a ranking. It is a role-based comparison for regulated buyers. Solution Best for Main strengths Limitations Microsoft Purview Microsoft-centric enterprises needing strong data security, compliance, audit and catalog foundations Strong fit for AI-powered data governance solutions, including data governance, audit, information protection, compliance and lifecycle management Less of a dedicated standalone AI risk suite; works best as a control foundation inside a broader Microsoft architecture IBM watsonx.governance Large regulated enterprises needing policy-to-control mapping across hybrid environments Strong governance graph, policy mapping, continuous reporting, regulatory content and AI/GRC integration Can be heavyweight for organizations looking for a narrow or lightweight use case Google Gemini Enterprise Agent Platform Google Cloud users building models and agents inside one engineering stack Strong model evaluation, registry, monitoring, secure development and governed enterprise-agent capabilities More platform-centric than governance-program-centric; may require additional compliance orchestration Credo AI Enterprises wanting centralized AI inventory, risk intelligence and regulatory mapping Strong registry, shadow-AI discovery, policy packs, evidence generation and governance across models, agents and applications Some teams may still pair it with separate model platforms or observability tools Dataiku Govern Organizations wanting governance embedded into the AI delivery workflow Strong workflows, registries, sign-off rules, audit timelines, LLM registry and growing agent-management capabilities Best fit when Dataiku is already part of the AI operating model Fiddler AI Runtime-heavy environments focused on monitoring, guardrails and observability Strong for continuous evaluation, root-cause visibility, inline enforcement and agentic monitoring More specialized around observability and runtime control than full enterprise management-system governance Arthur AI Teams prioritizing agent discovery, evaluation, observability and guardrails Good coverage of agent discovery, performance evaluation, built-in guardrails and model-agnostic support Less public emphasis on regulatory content libraries and formal enterprise compliance workflows MLflow Engineering-led teams needing open-source observability, evaluations, registries and model management Useful open-source backbone for custom AI governance stacks Not an out-of-the-box regulatory governance suite Evidently Teams needing open-source testing, monitoring and dashboards Strong for evaluating, testing and monitoring ML and LLM systems Not a complete governance operating system for policy, accountability or regulatory workflows Giskard LLM and agent teams focused on testing, red-teaming and evaluation Useful for LLM and agent safety, security and validation workflows Not a full enterprise governance suite with broad policy packs and formal approval routing AIF360 / Fairlearn Organizations needing open-source fairness assessment and bias mitigation Mature tooling for detecting and mitigating bias Best treated as components inside a wider governance design, not as end-to-end solutions for AI governance The practical pattern is clear. Platforms such as IBM, Credo AI and Dataiku are closer to end-to-end governance layers. Microsoft Purview and Google’s platform are powerful when governance is tightly linked to data estates and cloud engineering. Fiddler and Arthur are strongest where runtime performance, decision lineage, agent control and guardrails matter most. Open-source projects are indispensable for cost-effective experimentation and specialized controls, but they usually need architectural composition before they resemble full enterprise AI governance solutions. 11. Open-Source vs Commercial AI Governance Tools Organizations considering the best open-source AI governance solutions 2026 should take a toolkit view rather than look for one universal platform. Open-source is strong in technical subdomains: fairness and bias mitigation with AIF360 and Fairlearn, observability and drift monitoring with Evidently, evaluation and testing for LLM agents with Giskard, and AI engineering workflows with MLflow. These tools can be highly valuable, especially for engineering-led organizations. However, they are usually not full business governance systems. They do not, by themselves, deliver the full mix of regulatory mapping, approval workflows, ownership assignment, cross-functional reporting and audit-ready evidence that commercial governance suites emphasize. Commercial tools, by contrast, usually win on speed to governance. They package inventory, workflows, policy libraries, integrations, alerts, evidence capture and executive reporting in ways that better serve compliance, risk, procurement and audit teams. For regulated enterprises, the right answer is often hybrid: commercial governance platforms for enterprise control and reporting, supported by open-source tools for specific technical evaluations, monitoring or fairness checks. 13. Why Agentic AI Needs Separate Governance AI agents introduce a new governance challenge. Unlike traditional AI models that generate an output for a human to review, agents can plan, call tools, access systems, trigger workflows and perform multi-step actions. This changes the risk profile. Enterprises need enterprise AI agent governance solutions that can define what an agent is allowed to do, which systems it can access, what data it can use, when a human must approve an action and how every step is logged. Governance must cover the agent’s role, permissions, model behaviour, tool access, output quality, runtime monitoring and intervention paths. This is why agent governance should not be treated as a footnote to model governance. It requires its own inventory, approval workflows, control design, monitoring and incident response model. AI Agent Governance Checklist Every enterprise deploying AI agents should be able to answer these questions before production. ✓ What systems can it access? ✓ What data is the agent allowed to access? ✓ What actions is the agent allowed to perform? ✓ When is human approval required? ✓ Is every action logged? ✓ Can the agent be stopped immediately? ✓ Who is accountable for the agent? Organizations that cannot answer these questions before deployment will struggle to demonstrate effective governance once AI agents begin interacting with enterprise systems and business processes. 14. How to Choose the Right AI Governance Solution The best buying logic for regulated enterprises starts with the problem, not the vendor demo. If the main challenge is data sprawl, sensitive information control, audit and compliance across Microsoft environments, Microsoft Purview may be a strong foundation. If the priority is enterprise-wide policy management and regulatory mapping, IBM watsonx.governance, Credo AI or Dataiku Govern may be more relevant. If the business needs runtime quality control, observability, guardrails and agent monitoring, Fiddler AI or Arthur AI may become stronger candidates. If the organization is engineering-heavy and prepared to design its own operating model, open-source stacks based on MLflow, Evidently, Giskard and fairness libraries can be powerful. Second, test the platform against the regulatory footprint, not only the presentation. Regulated buyers should ask whether the solution supports risk classification, data quality and provenance, audit evidence, human oversight, third-party governance and post-deployment monitoring. Third, check whether the platform can support governance across the full AI estate: models, applications, agents, vendors, data pipelines and business processes. AI governance that only works for one model or one team will not scale across a regulated enterprise. 15. Why AI Governance Is More Than Software AI governance software can support discovery, workflows, evidence and monitoring, but it cannot define accountability on its own. Regulated organizations need a governance operating model that connects business owners, compliance, legal, data teams, security, IT, procurement and executive leadership. This is where AI governance consulting & solutions become essential. The platform is only one part of the answer. Organizations also need to define what AI use cases are allowed, how risks are classified, who approves deployment, what evidence is required, how vendors are assessed, how incidents are handled and how governance evolves as AI systems change. Without this operating model, even a strong platform becomes another dashboard. With the right governance framework, AI can move from pilots to production in a way that is controlled, auditable and aligned with business goals. 16. TTMS Project Insight: Governance Starts Before the Model One lesson we have seen repeatedly in client projects is that governance challenges rarely begin with the AI model itself. They usually start much earlier: with the quality of source documents, inconsistent business processes, fragmented knowledge and unclear ownership of information. In one TTMS project for a law firm, we developed an AI solution supporting court document analysis. While selecting the right language model was important, the biggest implementation effort focused on preparing trusted legal content, defining document workflows, validating AI-generated outputs and ensuring that lawyers remained in control of final decisions. Governance became an integral part of the solution rather than an additional compliance layer. The same pattern appears across regulated industries. Organizations often discover that successful AI adoption depends less on choosing the “best” model and more on establishing reliable governance around data, processes and human oversight from the very beginning. In our experience, organizations rarely struggle because they chose the wrong AI model. More often, they struggle because they underestimated the governance needed around it. Read more about this project in our AI implementation for court document analysis case study. You can also explore more examples in the TTMS case studies library. 17. How TTMS Helps Regulated Enterprises Govern AI TTMS supports organizations that need to move from AI ambition to governed AI implementation. As an AI consulting and strategy partner, TTMS helps regulated enterprises assess AI risk, design governance frameworks, select suitable governance architecture and operationalize controls across data, models, applications, vendors and agents. The company’s approach is strengthened by its ISO/IEC 42001-certified AI Management System. TTMS states that this system governs both internal and external AI-related projects delivered under the TTMS brand. This matters because AI governance is not only a client advisory topic. It is also a way of working that must be reflected in project delivery, documentation, risk management and operational oversight. For organizations using third-party AI tools, this is especially important. Governance is still required even when the AI model is not built in-house. Enterprises need to understand how external tools use data, how outputs are reviewed, what risks are introduced, which controls are required and how accountability is maintained. TTMS helps clients approach AI governance as a practical implementation challenge rather than a documentation exercise. The goal is not to slow innovation down, but to make AI adoption safer, more scalable and easier to defend in regulated environments. 18. From AI Governance Strategy to Practical Business Solutions Choosing the right AI governance platform is only one part of building a successful AI strategy. Organizations also need practical governance frameworks, clear policies, evidence workflows, vendor assessment, risk classification and implementation expertise that connects technology with business and regulatory requirements. At TTMS, we combine AI governance consulting & solutions with the development of secure, enterprise-ready AI products. Rather than offering a single generic AI platform, TTMS develops specialized solutions for individual business processes, allowing organizations to combine practical AI adoption with governance, security and regulatory compliance. This approach helps enterprises move from strategy to implementation: from selecting enterprise AI governance solutions and defining controls to deploying AI tools that support real operational needs in legal, document analysis, e-learning, knowledge management, localisation, AML, recruitment and software testing. AI4Legal helps legal teams analyse court documents, generate contracts and process hearing transcripts while maintaining full control over sensitive legal information. AI4Content enables secure document analysis and knowledge extraction, generating structured summaries and reports in controlled cloud or on-premise environments. AI4E-learning transforms internal documentation into complete e-learning courses, helping organizations scale AI literacy and workforce development. AI4Knowledge provides employees with governed access to organizational knowledge, procedures and internal documentation through conversational AI. AI4Localisation automates multilingual content translation while preserving terminology consistency and industry-specific language. AML Track supports anti-money laundering processes through automated screening, reporting and fully auditable compliance workflows. AI4Hire assists HR teams with CV analysis, candidate matching and resource allocation using transparent,>QATANA improves software quality by automating test management and AI-assisted test case generation in secure enterprise environments. All of these solutions are developed and delivered within TTMS’s AI Management System aligned with ISO/IEC 42001. This means clients benefit not only from innovative AI technology but also from established governance practices covering risk management, documentation, human oversight, security and regulatory compliance throughout the entire AI lifecycle. Whether your organization is evaluating enterprise AI governance solutions, looking for AI governance consulting & solutions, or planning to deploy AI in a regulated environment, TTMS helps turn governance into a practical business capability that enables innovation instead of slowing it down. FAQ What are the best AI governance solutions? There is no single universal winner. The best AI governance solutions depend on the enterprise problem. IBM watsonx.governance, Credo AI and Dataiku Govern are among the strongest broad governance suites. Microsoft Purview is highly relevant when data governance, compliance and Microsoft-stack integration dominate. Google’s Gemini Enterprise Agent Platform is strong for teams building governed agents and models in Google Cloud. Fiddler AI and Arthur AI can be excellent where runtime observability, agent control and guardrails are the priority. Open-source stacks can also be valuable, but usually as components rather than complete enterprise governance systems. What are the best open-source AI governance solutions in 2026? For buyers asking about the best open-source AI governance solutions 2026, the strongest answer is a toolkit view. MLflow is a broad open-source AI engineering base. Evidently is strong in testing and monitoring. Giskard is especially relevant for LLM and agent evaluation. AIF360 and Fairlearn are useful for fairness analysis and bias mitigation. However, most regulated enterprises will still need additional workflow, policy, reporting and audit layers on top. Can AI governance be automated? Yes, but only partially. Inventory, control mapping, evidence collection, recurring checks, continuous evaluations, alerts and parts of reporting can be automated effectively. Accountability decisions, material risk acceptance, exceptions and final approvals should remain under human oversight. The best automated AI governance solutions support governance teams instead of replacing them. Do organizations need ISO/IEC 42001 if they only use third-party AI tools? Certification is not always mandatory, but the standard is highly relevant for organizations using AI in regulated, customer-facing, high-impact or procurement-sensitive contexts. ISO/IEC 42001 is designed for organizations providing or using AI-based products and services. Even companies relying on external AI tools still need oversight, documentation, vendor accountability, data controls, risk assessment and human review. How should enterprises govern agentic AI? Enterprises should treat AI agents as a higher-governance category than ordinary chatbots. Agents need inventory, role and permission boundaries, model evaluation, action controls, logging, runtime monitoring and intervention paths for unsafe or off-policy behaviour. This is why the market is shifting toward enterprise AI agent governance solutions and why agent governance should be designed separately from traditional model governance. What Do Analyst Ratings Say About AI Governance Solutions? Publicly available best AI governance solutions analyst ratings should be treated carefully because many detailed comparisons from Gartner, Forrester and IDC sit behind paywalls. Still, public vendor disclosures and analyst mentions show a clear direction of travel. The market is rewarding platforms that provide centralized AI inventory, risk management, continuous monitoring, policy enforcement, evidence generation and agent/runtime governance. This is also why the search intent behind best AI governance solutions risk management 2026 is shifting away from one-time ethics checklists and toward continuous control planes. For regulated enterprises, this is the right direction. AI governance is converging with operational resilience, cybersecurity, data governance and enterprise risk management.

Read
GPT-5.6 from OpenAI – What’s New? Pricing, Features, and Business Applications

GPT-5.6 from OpenAI – What’s New? Pricing, Features, and Business Applications

For now, we can only talk about GPT-5.6 in Europe with a mix of professional curiosity and a slight sense of envy. OpenAI has initially made GPT-5.6 available only to a small group of selected partners working with the U.S. administration to evaluate the model’s safety, including potential cybersecurity risks. That’s why we prepared this article as a structured analysis based on official OpenAI materials, technical documentation, early expert evaluations, and publicly available market information. In this article, you’ll learn: What has changed in GPT-5.6 compared to GPT-5.5 and earlier OpenAI models? How do Sol, Terra, and Luna differ, and when should you use each model? How does GPT-5.6 compare with Claude, Gemini, DeepSeek, Grok, and other leading AI models? Which business areas are likely to benefit the most from GPT-5.6? OpenAI’s official statement reads: “We do not believe this government access process should become the long-term standard. It prevents our best tools from reaching the users, developers, businesses, cybersecurity defenders, and global partners who need them.” OpenAI says that broader availability is expected in the coming weeks. We look forward to updating this introduction with our own hands-on experience as soon as GPT-5.6 becomes available more widely. 1. GPT-5.6 – The Biggest Changes Compared to Previous Models 1.1 A New GPT-5.6 Architecture – Three Models Instead of One Universal Model The biggest change is architectural rather than incremental. OpenAI is moving away from the idea of a single flagship model for every task and introducing a family of models with distinct capability levels. In the new naming scheme, the version number represents the generation, while Sol, Terra, and Luna identify individual models that can evolve independently. If OpenAI continues down this path, future releases may no longer follow a simple GPT-5.5 → GPT-5.6 → GPT-5.7 progression, but instead develop as parallel model families. First, an important clarification: Sol, Terra, and Luna are not “modes” in the strict sense. They are three separate models within the GPT-5.6 family. The publicly announced operating modes currently include max reasoning effort and ultra, both available for Sol. Before we discuss them, let’s first look at how the three GPT-5.6 models differ and how OpenAI positions each of them. Model Positioning Best Use Cases Official API Pricing What We Know for Certain GPT-5.6 Sol Flagship model Most demanding tasks: advanced analysis, software development, AI agents, cybersecurity, and complex projects USD 5 input / USD 30 output per 1M tokens Supports max reasoning effort and ultra; the most capable model in the family GPT-5.6 Terra Balanced model Everyday business work, document analysis, automation, and the best quality-to-cost ratio USD 2.50 / USD 15 According to OpenAI, delivers GPT-5.5-level performance at roughly half the API cost GPT-5.6 Luna Fastest and most affordable model High-volume workloads, large-scale automation, frontline assistants, and cost-sensitive tasks USD 1 / USD 6 The fastest and most cost-efficient model in the GPT-5.6 family OpenAI describes ultra as a mode that uses sub-agents to speed up complex tasks. In practice, this means GPT-5.6 performs much better when a task requires multiple steps rather than a single answer. It can analyse large software projects, use external tools, conduct in-depth research, help identify software bugs, organise technical analysis, and prepare structured action plans. For organisations, this means higher efficiency in complex business processes, but also a greater need for monitoring, logging, and access control. 1.2 Stronger Reasoning and AI Agents – What Are max Reasoning Effort and ultra? The second major change is how the model approaches difficult tasks. For Sol, OpenAI introduces a new max reasoning effort level, allowing the model to spend more time analysing a problem before generating an answer. It also introduces ultra, a mode designed for the most complex tasks. In this mode, the model can break work into smaller stages and analyse different parts of a problem in parallel, reaching a solution more efficiently. This is more than a simple interface update. It reflects OpenAI’s shift from treating AI as a system that answers questions to one that helps complete entire tasks. 1.3 Better Programming, Cybersecurity and Scientific Research The third major improvement focuses on software development and tool usage. GPT-5.6 Sol is positioned as a model built for complex programming tasks, especially those that involve planning work, analysing repositories, debugging, using terminal environments, and completing multiple steps rather than simply generating code snippets. OpenAI highlights its strong performance on Terminal-Bench 2.1, a benchmark measuring how well AI models handle realistic software engineering tasks, as well as GPT-5.6’s availability through the API and Codex. For development teams, this represents an important shift. Rather than serving only as a coding assistant, GPT-5.6 increasingly supports the entire software development lifecycle—from analysing problems and refactoring code to generating tests and assisting with CI/CD workflows. The greatest benefits are likely to be seen by teams working on large software projects where AI can help manage complexity. Cybersecurity and scientific research are another area where GPT-5.6 has improved. According to OpenAI’s safety documentation, Sol and Terra can help identify vulnerabilities in IT systems and analyse how they could potentially be exploited. At the same time, internal testing showed that the models were not able to carry out complete attacks against well-protected systems on their own, highlighting both their growing capabilities and their current limitations. OpenAI and independent evaluators also report strong performance in biology and cybersecurity benchmarks, showing that GPT-5.6 is evolving beyond software development into a tool for highly technical and specialised domains. 1.4 Better Analysis of Documents, Images and Complex Data Another major improvement is GPT-5.6’s ability to work with different types of information. Rather than being viewed simply as a text model, GPT-5.6 is increasingly becoming part of a broader system for working with documents, images, research materials and business data. In practice, this means it is better suited to tasks that require combining multiple sources of information, such as reports, presentations, screenshots, technical documentation, meeting notes and visual materials. Instead of simply summarising individual files, the model can compare information, identify relationships and help build meaningful conclusions from different data formats. This is also where the difference between a standalone language model and a complete business solution becomes most apparent. Analysing enterprise documents requires more than just generating answers—it also involves access control, trusted sources, reporting workflows and compliance with company data policies. At TTMS, this is exactly the kind of functionality we build into solutions such as AI4Content. 1.5 GPT-5.6 Is More Autonomous, but Also Requires More Oversight OpenAI makes it clear that greater autonomy must be matched by stronger human oversight. According to the company’s safety documentation, GPT-5.6 Sol is more persistent than its predecessor when trying to complete a user’s objective and may occasionally take actions that go beyond the user’s original intent, although such cases remain relatively rare. Independent experts have reached similar conclusions. METR (Model Evaluation & Threat Research), an independent organisation specialising in evaluating advanced AI systems, found that GPT-5.6 Sol was more determined to complete tasks in certain tests, even if that meant attempting to bypass the rules of the testing environment. Meanwhile, Apollo Research, which studies AI safety, found no evidence that GPT-5.6 is more likely than previous models to take undesirable autonomous actions. In practice, this means GPT-5.6 can be more effective in long-running, agentic tasks, but it should operate within a well-designed environment that includes activity logging, access controls, human review and appropriate governance. 1.6 GPT-5.6 Features OpenAI’s Most Advanced Safety Architecture Yet OpenAI presents GPT-5.6 not only as a more capable model, but also as one designed for safer enterprise deployment. The model is intended to recognise risky prompts more effectively, reduce opportunities for misuse and operate within environments that provide stronger control over access, monitoring and usage policies. In practice, this means multiple layers of protection. Some safeguards are built directly into the model, others operate while responses are being generated, and others monitor suspicious usage patterns. Imagine a user repeatedly asking similar questions in slightly different ways to bypass the model’s safeguards and obtain instructions they should not receive. If the system detects a high risk of misuse, it can refuse the request, apply additional safeguards or route the interaction through stricter security controls. OpenAI also applies different access levels and extensive automated safety testing designed to determine whether GPT-5.6 can be manipulated into breaking its own safety rules—for example through jailbreak attempts. According to the company, these automated evaluations consumed more than 700,000 A100-equivalent GPU hours. This does not mean GPT-5.6 is immune to mistakes or misuse, but it does show that security has become a dedicated product layer rather than simply another part of model training. 1.7 GPT-5.6: Greater Flexibility and Lower AI Deployment Costs From a business perspective, one of the biggest changes is that organisations no longer need to rely on the most powerful—and most expensive—model for every task. Sol can be reserved for expert analysis, AI agents and technically demanding projects, while many day-to-day processes can run on the more affordable Terra or Luna models. This changes the economics of AI adoption. Organisations can now match the cost of a model to the value of the task, using different models for strategic analysis, high-volume customer interactions, document automation or internal business support. 2. How to Choose the Right GPT-5.6 Model and Mode for Your Task Using GPT-5.6 follows a simple process. First, you choose one of the three models: Luna, Terra or Sol. If you select Sol, you can also choose between two additional operating modes: max reasoning and ultra. Deep Research works independently of the selected model and is designed for comprehensive investigations across multiple sources, helping organise, analyse and synthesise information into coherent conclusions. Task Luna Terra Sol Max reasoning Ultra Deep Research Why This Choice? Fast responses and chatbots ✅ – – Lowest cost and very fast responses. Document classification ✅ ✅ – – Usually does not require advanced reasoning. Marketing content creation ✅ – – A good balance between quality, speed and cost. Legal contract and document analysis ✅ ✅ Complex documents benefit from deeper reasoning. Financial analysis and reporting ✅ ✅ Accuracy, consistency and stronger reasoning are essential. Programming and code review ✅ ✅ Additional reasoning time improves coding quality. Refactoring large software projects ✅ ✅ Ultra performs better in complex, multi-stage development tasks. Complex agentic workflows ✅ ✅ Ultra uses sub-agents to handle sophisticated workflows. Preparing reports from multiple sources ✅ ✅ Deep Research searches, compares and analyses multiple sources automatically. Expert articles and market analysis ✅ ✅ ✅ Combines in-depth research with advanced reasoning for the highest-quality results. Combining in-depth research with strong reasoning quality produces the best results. In practice, GPT-5.6 should not be treated as one model for every task, but as a set of configurations that can be matched to the difficulty of the task, the expected quality of the output, and the depth of research required. 3. What Will GPT-5.6 Pricing Look Like? The API pricing for the GPT-5.6 family is structured as follows: Sol – USD 5 / USD 30 per 1M input/output tokens, Terra – USD 2.50 / USD 15, Luna – USD 1 / USD 6. Sol remains at the same pricing level as GPT-5.5, so there is no price jump for the flagship model class. What is interesting is that OpenAI is clearly creating more affordable entry points: Terra is positioned as offering performance competitive with GPT-5.5 at roughly half the cost, while Luna is clearly focused on the best balance between quality and price. 4. The Evolution of OpenAI Models GPT-5.6 is best understood in a broader context. It is not just another model release with better benchmark results. It shows a shift in how OpenAI designs AI systems: from one universal model to a family of models with different costs, capabilities and use cases. Generation Release Parameters / Architecture, if Disclosed Context Length Multimodality Key Improvement Typical Business Use Cases GPT-1 2018 12-layer decoder-only Transformer, 768 hidden size, 12 attention heads 512 tokens No Generative pre-training as a universal transfer learning foundation Classification, basic NLP, research experiments GPT-2 2019 Up to 1.5B parameters; four variants from 117M to 1.542B 1,024 tokens No Major improvement in text generation and zero-shot transfer Content generation, summaries, experimental copywriting GPT-3 2020 175B parameters Not fully specified in the launch materials No Few-shot learning at production scale Chatbots, text automation, AI prototypes GPT-3.5 2022 Model from the GPT-3.5 series, fine-tuned for dialogue Later GPT-3.5 Turbo API versions supported 16k by default No Commercialisation of high-quality conversational AI through ChatGPT Support, FAQs, internal assistants, first enterprise deployments GPT-4 2023 Architecture and size not disclosed; large-scale multimodal model Not fully specified in the technical launch report Yes, image and text input Major leap in reasoning, exam performance, instruction following and safety Document analysis, expert knowledge work, advisory tasks, high-stakes deployments GPT-4o 2024 Frontier model optimised for practical multimodality Not explicitly stated on the cited launch page Yes, text, image, voice and broader product-level multimodality Omni model: faster, cheaper and more natural multimodal interaction Voice assistants, image analysis, customer service, multimodal copilots GPT-5 2025 Unified system with routing between fast and deeper reasoning paths 400k, with up to 128k output in API documentation Text and image input, text output Automatic routing, higher usefulness, fewer hallucinations and better tool use AI agents, software development, knowledge work, expert analysis GPT-5.5 2026 Frontier model for complex work; later matched by Sol-level pricing in GPT-5.6 1M Strongly oriented around documents and tools in ChatGPT and API Better persistence in long-running tasks, software work, research and data analysis Research, document analysis, modelling, customer operations, finance GPT-5.6 2026 No full public parameter specification; Sol/Terra/Luna model family Not publicly disclosed in a separate preview model card Recent OpenAI models support text and image input, but GPT-5.6 preview does not yet have a full public specification card Capability tiers, max reasoning, ultra mode, sub-agents and a stronger deployment safety layer Agentic software workflows, cybersecurity, enterprise document work, high-volume automation with better cost control The shortest way to summarise this evolution is this: from GPT-1 to GPT-3, OpenAI mainly scaled the model itself; from GPT-3.5 to GPT-4, it refined the human-model interface; and from GPT-5 onwards, it has been building a broader AI work system with routing, tools, longer task horizons, cost control and stronger safety layers. GPT-5.6 shows this direction clearly: OpenAI is moving from standalone chatbots towards systems that support work, automation and decision-making. 5. GPT-5.6 in Business: Where Will Companies Feel the Biggest Change? 5.1 GPT-5.6 in Marketing – Faster Content Operations and Better Data Analysis In marketing, the biggest change is about scale and cost efficiency in working with content and data. Sol can be used for research, strategy, more difficult analyses and multi-variant campaigns, while Terra and Luna are better suited to high-volume tasks: paraphrasing, content tagging, creative drafts, summaries, extracting insights from research and automating everyday content operations. In similar scenarios, AI4Localisation can be a strong fit. It is a TTMS solution supporting translation and localisation of business content. With AI, organisations can prepare multilingual materials faster while maintaining consistent terminology and communication style. 5.2 GPT-5.6 for Developers – Code Review, Refactoring and AI Agents The change is especially visible in software development. GPT-5.6 Sol is expected to perform better in long, multi-step tasks such as repository analysis, bug detection, refactoring, test generation and support for work in environments such as the API or Codex. This means AI can help not only with writing individual code snippets, but also with organising larger development tasks. This does not mean engineering oversight can be removed. The more a model can do independently, the more important code review, testing, permission limits and clear rules become. Teams need to decide what AI can execute automatically and what still requires human approval. 5.3 GPT-5.6 in Customer Service – Ticket Automation and Consultant Support In customer service, Terra and Luna may be especially useful as faster and more affordable GPT-5.6 variants. OpenAI positions Terra as a model for everyday business tasks, while Luna is the fastest and cheapest option in the family. This fits well with first-line support work: organising tickets, assigning priority, preparing response drafts, extracting key information from customer requests and suggesting next steps to consultants. 5.4 GPT-5.6 in HR and Recruitment – CV Analysis, Onboarding and Recruiter Support In HR, the greatest value of GPT-5.6 may come from combining better information analysis with more flexible usage costs. In practice, this means support with summarising CVs, comparing candidates, organising recruitment notes, preparing shortlists and creating onboarding plans. Terra may often be more cost-effective than Sol here, because many recruitment tasks are performed at scale but do not require the most advanced level of reasoning. In this area, AI4Hire fits naturally as a TTMS tool for CV analysis and matching skills to projects. It automates profile assessment, generates recommendations and helps teams find people who best match a specific requirement faster. 5.5 GPT-5.6 in Compliance – Document Analysis and Regulatory Support In compliance, accuracy, consistency and alignment with procedures matter most. GPT-5.6 may be useful here because OpenAI highlights several safety layers: response monitoring during generation, detection of suspicious usage patterns and different levels of model access. This does not mean GPT-5.6 can make regulatory decisions on its own. It can, however, support policy analysis, document review, preparation of evidence materials, checking whether outputs follow internal procedures and internal audits. AI4Legal uses similar capabilities in the legal sector. It is a TTMS solution supporting law firms in document analysis, contract preparation, work with case files and transcript processing. In practice, it shows that the biggest value of models such as GPT-5.6 comes not from giving users access to the model itself, but from integrating AI into a specific business process. Another example of AI in compliance is AML Track, a TTMS solution supporting AML processes such as customer verification, sanctions list screening, report preparation and audit trail maintenance. It shows that in compliance, AI does not need to replace expert judgement. It can organise data, automate repetitive work and support alignment with regulatory requirements. 5.6 GPT-5.6 in Finance – Report Analysis, Due Diligence and Controlling Support In finance and controlling, the real value of GPT-5.6 is likely to appear where teams need to combine documents, calculations, multi-step analysis and repeatability. GPT-5.5 was already positioned as a model that performs well in data analysis, information retrieval and work with large document sets. With GPT-5.6, organisations can more easily match the cost of AI usage to a specific task while gaining more advanced agentic capabilities. The biggest impact will therefore be felt not by simple financial chatbots, but by teams working with large volumes of documents and data: due diligence, report analysis, KYC processes, extracting key metrics and preparing materials for decision-makers. For now, these are conclusions based on the capabilities described by OpenAI and early tests, not yet on widely documented GPT-5.6 finance deployments. 5.7 GPT-5.6 in E-learning – Faster Training Creation and Personalised Learning In e-learning, GPT-5.6 may offer very practical benefits: faster breakdown of large knowledge sets into modules, creation of assessment questions, transformation of documents into training formats, personalisation of learning paths and the development of internal tutors. If this cost-and-capability model split continues, Terra and Luna may be used for high-volume content production and updates, while Sol can support the design of more advanced, expert-level or highly contextual materials. This is also the direction behind AI4E-learning, a TTMS tool that helps turn company materials, documents and presentations into ready-to-edit e-learning courses that can be exported to LMS platforms. 5.8 GPT-5.6 in Software Testing – QA Support and Test Automation GPT-5.6 may also be especially useful for QA teams. The model can help generate test cases, analyse regression issues, interpret logs, recreate error paths and prepare drafts of automated tests. What also matters is that companies can choose the model variant based on the task: Sol for more complex troubleshooting, Luna for large volumes of simpler, routine testing tasks. QATANA follows this direction as well. It is a TTMS solution for AI-supported software test management, helping QA teams generate test cases, analyse requirements, organise the testing process and improve control over application quality. 6. Is GPT-5.6 the Best LLM Today? A Comparison with Competitors Area Is GPT-5.6 the Best Here? Main Competitor Programming ✅ Yes Claude Opus AI Agents ✅ Yes Claude Documents ✅ Yes Claude Multimodality ⚠️ Tie Gemini Price ❌ No DeepSeek On-premise ❌ No Mistral / Llama Google Workspace ❌ No Gemini 6.1 Programming – GPT-5.6 Sol or Claude Opus? Both models are currently among the strongest options for software development. Claude Opus has long been valued for its ability to work with large code repositories and analyse existing projects. GPT-5.6 Sol, however, appears to go a step further thanks to its agentic capabilities, Max reasoning and Ultra modes, and strong results in benchmarks such as Terminal-Bench 2.1. If a task requires not only writing code, but also planning, using tools and completing several stages of work, GPT-5.6 Sol is likely to have the advantage. 6.2 AI Agents – Where OpenAI Has a Clear Advantage This is currently one of GPT-5.6’s strongest areas. OpenAI is developing the model not only as a classic chatbot, but as a platform for AI agents that can plan actions, use tools and carry out complex tasks. Claude is also developing agentic capabilities, but it does not currently offer a direct equivalent of Ultra, which uses sub-agents to solve complex problems in parallel. 6.3 Document Analysis – GPT-5.6 or Claude? Claude has long been considered one of the best models for working with long documents and complex text. GPT-5.6 Sol appears to be very close in terms of document analysis quality, while its stronger reasoning may help it draw conclusions from multiple sources at once. In practice, both models are likely to perform at a very high level, although GPT-5.6 offers broader options for using document analysis inside agentic business processes. 6.4 Multimodality – Gemini Still Sets the Direction If the main task is to analyse text, images, video and audio together, Gemini remains a very strong option. This is mainly because it was designed from the beginning as a natively multimodal model and is deeply integrated with Google’s ecosystem. GPT-5.6 also performs well in multimodal tasks, but in this area it is difficult to name a clear winner. 6.5 Price – DeepSeek Remains Hard to Beat When it comes to API costs, DeepSeek still clearly undercuts most major competitors. For organisations handling millions of requests per month, the price difference can translate into substantial savings. The trade-off is lower transparency around safety and a weaker tool ecosystem compared with OpenAI. 6.6 Local Deployments – Where Mistral and Llama Have the Advantage Not every organisation can use models that run only in the cloud. Companies in finance, public administration or defence often need full control over infrastructure and data. In such cases, models that can be run on private servers, without sending data to an external cloud, have an advantage. Examples include Mistral Large 3 and Llama 4. 6.7 Google Workspace – Gemini’s Natural Environment Organisations that use Gmail, Google Docs, Google Drive or Google Meet every day will often gain the most from Gemini. The model was designed for close integration with Google’s services, which allows it to use data from that ecosystem and support everyday user workflows. There is no single AI model today that clearly wins in every category. GPT-5.6 Sol appears to be one of the most versatile options for business use, but the best model still depends on the use case, budget, security requirements and the environment in which it will be used. 7. What Does GPT-5.6 Mean for Companies? GPT-5.6 does not look like a routine model update. More important than better answer quality is the fact that OpenAI gives companies more choice: Sol for difficult tasks, Terra for everyday work and Luna for processes where scale and cost matter most. For businesses, this means one thing: access to GPT-5.6 alone will not be enough. The real value will come from placing the model inside a specific process, connecting it with organisational knowledge, securing the data and clearly defining where AI supports people and where people still make the final decision. Full GPT-5.6 availability in Europe may still take some time, but the direction is already clear. The companies that benefit most will not simply be those that adopt the newest model first, but those that match AI to real tasks, costs, data and security rules. If you are considering how to introduce AI into your organisation, explore our AI Solutions or contact our team to discuss which approach fits your business processes best. Is GPT-5.6 available in Europe? Not yet for general public use. While ChatGPT and the OpenAI API are available across most European countries, GPT-5.6 has so far been released through a limited preview programme for a small group of trusted partners. This rollout is not specific to Europe – it affects nearly all markets outside the preview programme. OpenAI has confirmed that broader availability will be introduced gradually. When will GPT-5.6 become available in Europe? OpenAI has not announced a specific launch date for Europe. The company has stated that wider access is expected in the coming weeks, with availability expanding progressively across ChatGPT, the API and other OpenAI products. As with previous major releases, the rollout is likely to happen in stages rather than all at once. Are Sol, Terra and Luna GPT operating modes? No. Sol, Terra and Luna are three separate models within the GPT-5.6 family, not operating modes. The actual operating modes currently described by OpenAI are max reasoning effort and Ultra, both available for GPT-5.6 Sol. Each model is designed for different performance, cost and business scenarios. What is GPT-5.6 Sol? GPT-5.6 Sol is the flagship model in the GPT-5.6 family. It is designed for the most demanding tasks, including advanced reasoning, software development, AI agents, cybersecurity and complex enterprise workflows. Sol also supports the max reasoning effort and Ultra modes, making it the most capable model in the family. What is GPT-5.6 Terra? GPT-5.6 Terra is the balanced model in the GPT-5.6 lineup. OpenAI positions it as the best choice for everyday business work, document analysis and automation tasks where organisations need strong performance without paying for the most advanced model. According to OpenAI, Terra delivers performance comparable to GPT-5.5 at roughly half the API cost. What is GPT-5.6 Luna? GPT-5.6 Luna is the fastest and most affordable model in the family. It is intended for high-volume workloads such as chatbots, customer support, document classification and large-scale business automation. Luna is designed for situations where response speed and cost efficiency matter more than maximum reasoning capability. What does max reasoning effort mean in GPT-5.6? Max reasoning effort is an optional operating mode available for GPT-5.6 Sol. Instead of generating an answer as quickly as possible, the model spends more time analysing the problem before responding. This often improves performance in complex reasoning, programming, research and analytical tasks where accuracy is more important than speed. What is Ultra mode in GPT-5.6? Ultra is the most advanced operating mode available for GPT-5.6 Sol. OpenAI describes it as a mode that uses sub-agents to tackle complex problems by breaking them into smaller tasks and processing them in parallel. It is designed for long, multi-step workflows rather than simple question answering. How much does GPT-5.6 cost through the API? According to OpenAI’s published API pricing: GPT-5.6 Sol: USD 5 input / USD 30 output per one million tokens GPT-5.6 Terra: USD 2.50 input / USD 15 output GPT-5.6 Luna: USD 1 input / USD 6 output These pricing tiers allow organisations to choose the model that best matches both the complexity of the task and the available budget. Will GPT-5.6 be available through the API? Yes. OpenAI has confirmed that GPT-5.6 is being rolled out through the API as part of the preview programme and will become more broadly available as the rollout expands. The company also plans to make the models available across ChatGPT, Codex and other OpenAI services. Is GPT-5.6 safer than previous OpenAI models? OpenAI describes GPT-5.6 as its most security-focused model family to date. It introduces multiple layers of protection, including safeguards built into the model, real-time safety monitoring, usage pattern detection and different access levels. Independent researchers have not found evidence that GPT-5.6 is more likely than previous models to engage in undesirable autonomous behaviour, although its greater capabilities also make proper governance and human oversight more important. Is GPT-5.6 better suited for business than GPT-5.5? For many organisations, yes. GPT-5.6 introduces three specialised models instead of relying on a single universal model, allowing businesses to balance performance and cost more effectively. Companies can reserve Sol for highly complex work while using Terra or Luna for everyday automation, making enterprise AI deployments more flexible and cost-efficient than before. How can I get access to GPT-5.6? At the moment, access is limited to organisations participating in OpenAI’s preview programme. For everyone else, the best option is to wait for the wider rollout that OpenAI has announced for ChatGPT, the API and its other products. Availability is expected to expand gradually rather than becoming available worldwide on a single release date.

Read
Why Semantic Layers Matter for Enterprise AI

Why Semantic Layers Matter for Enterprise AI

Many companies are already using AI in one way or another. Employees ask AI assistants for help, business teams test copilots, and technology leaders look at agentic AI as a way to automate more complex processes. But after the first wave of experiments, one thing is becoming clear: a powerful language model is not enough to create real business value. The problem is usually not the model, but the context around it. AI can only give useful answers when it understands what company data actually means. That means knowing how metrics are defined, which sources can be trusted, how different systems relate to each other, and what rules apply to specific business processes. This is why more organizations are now investing in semantic layers, data governance, and architectures that make enterprise data understandable not only for people, but also for AI systems. 1. Why Enterprise AI Often Produces Inconsistent Answers Modern large language models are remarkably capable. They can summarize information, generate reports, answer questions, and assist with decision-making. However, they do not inherently understand how a specific organization operates. Consider a seemingly simple question: “What is our current customer profitability?” To answer accurately, an AI system may need to access data from ERP systems, CRM platforms, financial applications, customer support tools, and analytics environments. Even if all required data is available, another challenge emerges. Different departments may use different definitions for the same metric. Finance may calculate profitability differently than sales. Operations may classify customers differently than marketing. Regional teams may follow different reporting standards than global headquarters. When AI interacts with fragmented or inconsistent information, it can produce answers that appear correct while actually reflecting conflicting business assumptions. This is one of the primary reasons many organizations struggle to scale AI beyond isolated use cases. 2. The Missing Piece: Business Context For years, organizations focused on collecting and storing data. Data lakes, warehouses, analytics platforms, and reporting tools became central elements of enterprise architecture. AI introduces a new requirement. Systems must not only access data – they must understand it. A customer record, product identifier, invoice number, or performance metric may seem straightforward from a technical perspective. However, every organization has its own definitions, relationships, policies, and business rules that shape how information should be interpreted. Without this context, AI is forced to infer meaning from technical structures alone. With context, AI can generate answers that align with how the business actually operates. This distinction becomes even more important as organizations move from AI assistants toward AI agents capable of making recommendations, triggering workflows, and supporting operational decisions. 3. What Is a Semantic Layer? A semantic layer provides a business-friendly interpretation of enterprise data. Instead of exposing raw tables, schemas, and technical metadata, it creates a structured representation of business concepts that both humans and AI systems can understand. A semantic layer typically defines: Business metrics and KPIs Relationships between datasets Common terminology Calculation logic Data ownership Business rules and policies For example, when an executive asks about revenue, customer churn, inventory availability, or working capital, the semantic layer ensures that everyone – including AI systems – uses the same definitions. This creates a single source of business truth that can be reused across reporting, analytics, applications, and AI initiatives. 4. Why Semantic Layers Matter for AI The value of a semantic layer increases dramatically in AI-driven environments. Traditional dashboards require users to interpret data manually. AI systems, on the other hand, must interpret information autonomously. Without a semantic layer, AI models may: Misinterpret business terminology Combine incompatible datasets Apply inconsistent KPI definitions Generate conflicting recommendations Reduce trust among business users With a semantic layer, AI systems gain access to the organizational context required to produce more accurate and consistent outputs. This is increasingly important for natural language interfaces, AI copilots, knowledge assistants, and agentic AI architectures. The quality of AI responses becomes directly tied to the quality of the semantic framework supporting them. 5. Why Governance Is Becoming a Strategic Requirement As AI becomes embedded in business processes, governance is evolving from a compliance topic into a strategic business capability. Organizations need confidence that AI systems are operating within defined boundaries and using trusted information. The growing importance of governance in the field of artificial intelligence is also reflected in emerging international standards. An increasing number of organizations are adopting the requirements of ISO/IEC 42001 – the world’s first standard defining the principles for establishing and maintaining an Artificial Intelligence Management System (AIMS). TTMS is among the pioneers of this approach, having achieved ISO/IEC 42001 certification as one of the first companies in Europe and the first organization in Poland. The certification confirms that our processes for designing, implementing, and managing AI solutions are aligned with internationally recognized standards for security, transparency, accountability, and risk management. Strong governance helps answer critical questions: Which datasets are approved for AI use? Who owns specific business definitions? How should sensitive information be protected? Which users can access which data? How can AI-generated outputs be monitored and audited? Without governance, AI may generate answers that are technically correct but operationally risky. With governance, organizations can scale AI adoption while maintaining trust, security, compliance, and accountability. This explains why governance is becoming a central component of modern enterprise AI platforms rather than an afterthought. 6. The Rise of Agentic AI Changes Everything The next phase of enterprise AI extends beyond answering questions. Organizations are increasingly exploring agentic AI systems that can execute tasks, coordinate workflows, analyze data, and support operational decisions. Unlike traditional AI assistants, these systems are expected to interact with business processes directly. That creates a much higher standard for data quality and contextual understanding. An AI agent responsible for inventory optimization, customer service, financial planning, or procurement cannot rely on ambiguous definitions or inconsistent information. It requires a governed environment where business concepts are clearly defined and consistently applied. This is precisely why semantic layers and governance frameworks are becoming foundational components of agentic architectures. 7. Why Leading Technology Vendors Are Investing in Semantic Architectures Across the enterprise software market, a common pattern is emerging. Technology providers are investing heavily in metadata management, business context layers, semantic models, governance capabilities, and AI-ready data architectures. The reason is straightforward. Organizations no longer need AI that can simply generate text. They need AI that understands their business. Enterprise value is created when AI can accurately interpret operational realities, financial metrics, customer relationships, regulatory requirements, and organizational objectives. The companies enabling this contextual understanding will play a critical role in the next generation of enterprise technology. 8. Building an AI-Ready Data Foundation For organizations evaluating their AI strategy, the focus should extend beyond model selection. Questions such as “Which LLM should we use?” remain important, but they are increasingly secondary to more fundamental considerations. Technology leaders should also ask: Do we have consistent KPI definitions? Can AI access trusted and governed data? Do business users agree on terminology? Is ownership of critical data clearly defined? Can our AI systems explain how conclusions were reached? The organizations that answer these questions successfully are more likely to achieve measurable business outcomes from AI investments. 9. Conclusion Enterprise AI is entering a new phase. The conversation is gradually shifting away from model performance and toward business understanding. Semantic layers, governance frameworks, and contextual data architectures are becoming critical enablers of trustworthy AI. They help transform disconnected data into business knowledge and ensure that AI systems operate with the context required to support meaningful decisions. As organizations move toward increasingly autonomous AI capabilities, competitive advantage will not come solely from access to advanced models. It will come from the ability to connect those models with trusted data, shared business definitions, and a clear understanding of how the organization operates. In the era of enterprise AI, context is becoming as important as intelligence itself. Organizations looking to move beyond AI experimentation should focus not only on selecting the right models, but also on building the data foundations that allow AI to deliver measurable business value. This requires a combination of data strategy, governance, system integration, and AI expertise. At TTMS, we help organizations design and implement AI solutions that connect advanced models with real business processes, enterprise data, and operational goals. Explore our AI solutions for business to learn how we can support your AI transformation journey. What are some signs that an organization needs a semantic layer? One of the clearest indicators is when different departments report different values for the same KPI. If finance, sales, and operations each have their own version of revenue, profitability, or customer metrics, AI systems will struggle to provide consistent answers. Other warning signs include long discussions about which data source is correct, difficulties scaling analytics initiatives, or a lack of trust in AI-generated insights. A semantic layer helps create a common business language that can be used across teams, applications, and AI solutions. Can small and mid-sized companies benefit from semantic layers, or are they only for large enterprises? While semantic layers are often associated with large organizations, smaller businesses can benefit as well. As companies adopt more systems and generate more data, maintaining consistency becomes increasingly difficult. Establishing common definitions and governance practices early can prevent future complexity and make AI initiatives easier to scale. For growing organizations, a semantic layer can serve as a foundation that supports expansion without creating data silos. How do semantic layers support regulatory compliance? Many industries operate under strict regulations regarding data access, privacy, reporting, and auditability. A semantic layer can help by ensuring that business metrics and definitions are applied consistently across the organization. When combined with governance controls, it becomes easier to track how information is used, who has access to it, and how AI-generated outputs are produced. This level of transparency can simplify compliance efforts and reduce operational risk. What is the relationship between semantic layers and knowledge management? Organizations often store valuable business knowledge in documents, spreadsheets, presentations, emails, and the experience of individual employees. A semantic layer helps connect structured business definitions with this broader organizational knowledge. As a result, AI systems can provide more relevant answers and recommendations by combining enterprise data with the business context behind it. This makes knowledge more accessible and less dependent on specific individuals. Will future AI systems automatically create semantic layers on their own? AI will likely play an increasingly important role in identifying relationships between datasets, suggesting business definitions, and helping build semantic models. However, organizations will still need human expertise to validate those definitions and ensure they reflect real business objectives. Business context, governance policies, and strategic priorities cannot be fully automated. Rather than replacing semantic layers, future AI systems will likely make them easier and faster to develop and maintain.

Read
1
239