15 ChatGPT Integrations with Business Apps in 2026

15 ChatGPT Integrations with Business Apps in 2026

How can ChatGPT integrations with business applications simplify everyday work in 2026? Here is a simple example: a client emails us asking for a project status update. At this point, we face half an hour of clicking between Google Drive, Slack, Asana and the CRM system. What if ChatGPT could collect information from all these sources in a single conversation and immediately prepare a summary, response or plan for the next steps? In this article: we examine 15 ChatGPT integrations with popular business applications that can make the scenario described above part of a company’s everyday workflow, we explain the differences between apps, integrations, plugins, GPTs and MCP servers, we present specific business use cases and highlight what should be checked before implementation, including permission scopes, data security, availability and costs. How do business application integrations extend ChatGPT’s capabilities? In the client enquiry scenario described above, the right set of integrations could work as follows: ChatGPT would find documents in Google Drive, summarise conversations in Slack, check task statuses in Asana and analyse the client’s data in the CRM system. Individual integrations may be available in ChatGPT as apps, connectors or MCP-based solutions. They make it possible to use data and selected functions from external services without leaving the conversation, and then prepare an up-to-date summary, a response for the client or a plan for the next steps. The available capabilities depend on the specific solution. Some integrations are more “passive” and are used mainly for searching and reading data. More “active” integrations support creating, updating and sending data. GPTs, apps, connectors, plugins and MCP – how do these concepts differ? The terminology surrounding ChatGPT extensions includes several related concepts. In our previous article, we described the ecosystem of the most useful ChatGPT plugins. In this comparison, we use the term “ChatGPT integrations” as an umbrella term for the different ways of connecting ChatGPT to business applications, their data and their functions. Concept Proposed definition ChatGPT integration An umbrella term for connecting ChatGPT to an external application, its data or its functions. An integration may be implemented as an app, connector, plugin, MCP server or GPT Action. App A function of an external service available directly in ChatGPT, sometimes with an interactive interface. Connector A ready-made connection that gives ChatGPT access to the data or functions of a specific service. Depending on the solution, it may support search, synchronisation or actions. MCP server A layer that gives ChatGPT access to selected tools, data and operations from an external system in accordance with the Model Context Protocol standard. Plugin An installable package that extends ChatGPT or Codex and may include instructions, skills, an MCP connection and an optional interface. These concepts describe different elements of the same ecosystem and are not always completely separate. Integration is the umbrella term for connecting ChatGPT to an external service. It may be available as an app, use a connector or MCP server, while a plugin may combine several of these elements into a ready-made workflow. How do you connect an app to ChatGPT step by step? Define the task the integration should perform. Open the app or plugin directory in ChatGPT. Select the appropriate service and start the connection process. Sign in to the external application and approve the required permissions. Open a new conversation and select the connected app. Test the integration using a limited dataset before deploying it across the entire team. How did we select 15 ChatGPT integrations with business applications? This comparison covers integrations that support recurring business processes and are available directly in ChatGPT or through documented MCP-based solutions. We considered five criteria: Frequency of use: the tool stores data or supports tasks performed by teams every day. Value of context: the connection gives ChatGPT access to information that significantly improves the quality of its output. Scope of actions: the integration supports searching, analysing, creating or updating data. Access control: the provider describes authentication, user permissions or administrative controls. Usefulness across multiple roles: the solution can support sales, marketing, operations, IT, product development or knowledge management. The availability of individual features depends on the ChatGPT plan, the external service plan, the country, workspace settings and administrator decisions. The catalogue and permission scopes should be checked immediately before implementation. 15 ChatGPT integrations with business applications in 2026 1. Google Drive integration with ChatGPT – searching and analysing company documents The Google Drive integration with ChatGPT enables users to work with materials stored in Drive, Docs, Sheets and Slides. Users can search for files, combine information from several documents, analyse spreadsheets and use existing materials as sources for a new report, brief or presentation. It provides the greatest value to teams with well-organised folders and consistent document naming conventions. ChatGPT can then locate the correct versions of proposals, reports, meeting notes and project materials more quickly. Best use case: preparing a project summary based on documents, a results spreadsheet and a status presentation. Example prompt: “Find materials in Google Drive related to Project X from the last 30 days and prepare a summary of decisions, risks and next steps.” The video shows how to connect Google Drive to ChatGPT, create an SEO-optimised blog post and save it as a document in Google Drive. It also highlights the importance of detailed prompts for improving the quality of generated content. 2. SharePoint integration with ChatGPT – access to organisational knowledge, procedures and files SharePoint is a natural source of information for organisations using Microsoft 365. It stores documents, intranet pages, procedures, policies and project materials. The SharePoint integration with ChatGPT enables users to find these resources and use them when preparing responses or documents. It is particularly useful in larger organisations where knowledge is distributed across sites, document libraries and teams. The SharePoint permission structure continues to determine which information is available to each employee. Best use case: finding current policies, instructions, templates and project documentation. Example prompt: “Based on the current procedures in SharePoint, prepare an onboarding checklist for a new supplier.” 3. Box integration with ChatGPT – secure analysis of company documents Box combines content management with access controls and is often used by organisations working with confidential documents. The Box integration with ChatGPT can retrieve data on demand or synchronise selected content. On-demand access retrieves the required information while a prompt is being processed, while synchronisation indexes approved resources in advance and speeds up searches across large repositories. The choice of access mode should take into account data classification, retention requirements and the expected response time. Best use case: analysing contracts, project materials, client documentation and approved company resources. Example prompt: “Find the current versions of documents for Client X in Box and identify discrepancies in the project scope.” 4. Gmail integration with ChatGPT – summarising correspondence and preparing replies The Gmail integration with ChatGPT enables users to search for messages, summarise long threads and prepare draft replies based on their email history. Gmail in ChatGPT is useful in sales, customer service, recruitment and day-to-day coordination when important decisions are distributed across multiple messages. To help the Gmail connector return an accurate result, specify the relevant period, senders, subject and expected outcome. ChatGPT can then find the appropriate messages and turn them into a summary, list of decisions or ready-to-use draft reply. Best use case: summarising an email thread, preparing a follow-up and identifying the commitments made by each party. Example prompt: “Summarise the correspondence with Company X from the last two weeks. List the agreed actions, deadlines and questions that still require a response.” The video shows how to connect Gmail to ChatGPT step by step using the official app. Once the Gmail integration with ChatGPT has been configured, users can search for messages, summarise long threads, find important information and prepare draft replies directly within the conversation. 5. Outlook Email integration with ChatGPT – analysing messages in Microsoft 365 The Outlook Email integration with ChatGPT enables users to find messages, analyse long email threads and prepare replies that take the conversation history into account. Outlook in ChatGPT is particularly useful for organisations using Microsoft 365. If ChatGPT is also connected to SharePoint and Microsoft Teams, it can combine email discussions with documents and team conversations. The Outlook Email integration operates only within sources approved by the organisation and available to the individual user. Best use case: preparing a client response based on email history and current project materials. Example prompt: “Find the latest email thread about renewing the contract with Company X and prepare a draft reply that addresses the outstanding issues.” 6. Slack integration with ChatGPT – summarising team conversations, decisions and actions The Slack integration with ChatGPT gives the model access to context from messages, files, channels and team member profiles. Slack in ChatGPT helps reconstruct the history of decisions, prepare project status updates and identify recurring problems in team conversations. The Slack MCP server also supports selected actions, such as sending messages and creating or viewing Canvas documents. The Slack integration with ChatGPT only uses channels available to the authenticated user and operates according to the rules configured by the administrator. Best use case: preparing a weekly status update covering decisions, blockers, owners and open questions. Example prompt: “Review the project channel from Monday onwards and prepare a status update covering completed actions, risks, decisions and tasks for the coming week.” 7. Microsoft Teams integration with ChatGPT – analysing conversations, meetings and tasks The Microsoft Teams integration with ChatGPT enables users to search and analyse messages from individual chats, group conversations and channels available to them. Microsoft Teams in ChatGPT can also work with Microsoft Planner plans and tasks. When the relevant actions are enabled, it can create chats and channels, as well as send messages and replies. On the Enterprise plan, the integration can also retrieve transcripts from scheduled meetings if the user has the appropriate permissions. Files shared in Teams channels are usually stored in SharePoint, so analysing them requires an additional connection between ChatGPT and SharePoint. Best use case: finding decisions in team conversations and turning them into summaries, tasks and status materials. Example prompt: “Review the conversations in the project channel from the last five days and prepare a list of decisions, open questions, responsible individuals and deadlines.” 8. Notion integration with ChatGPT – creating and updating company knowledge The Notion integration with ChatGPT enables users to read, create and update content on Notion pages directly from a conversation. Notion in ChatGPT can support product documentation, campaign plans, knowledge bases, feature specifications and implementation checklists. The Notion MCP server operates within the permissions of the signed-in user. A person with broad access to the workspace gives the integration an equally broad scope of data and operations, so it is worth beginning the implementation with clearly limited use cases and accounts with appropriately assigned roles. Best use case: transforming notes and analysis results into structured pages, databases and action plans. Example prompt: “Create a feature specification in Notion based on these notes. Add objectives, requirements, acceptance criteria, risks and open questions.” The video shows how to connect Notion to ChatGPT and work with content stored in a workspace. The Notion integration with ChatGPT enables users to search for information and create or update pages directly from a conversation. 9. Atlassian Rovo integration with ChatGPT – working with Jira, Confluence and Bitbucket The Atlassian Rovo integration with ChatGPT connects the model to Jira, Jira Service Management, Confluence and Bitbucket. Jira and Confluence content can be searched and summarised in ChatGPT, while users can also create and update tasks, tickets and pages using natural language commands. The Atlassian Rovo MCP server supports software development, ticket management, change management and documentation processes. OAuth 2.1 authentication preserves existing user roles and permissions, while actions affecting data should be subject to approval and monitoring. Best use case: creating tickets from meeting notes, updating statuses and connecting Confluence documentation with Jira tasks. Example prompt: “Based on this specification, create five Jira tasks with descriptions, acceptance criteria and priorities. Show me the proposed tasks before saving them.” 10. Asana integration with ChatGPT – creating tasks and managing projects The Asana integration with ChatGPT provides information about projects and portfolios, and allows users to create and assign tasks, set up new projects and monitor progress. Asana in ChatGPT can turn decisions made during a conversation into a structured plan saved directly in the work management system. The Asana integration with ChatGPT is useful for planning campaigns, implementations, product launches and cross-departmental initiatives. The integration produces the most accurate results when projects, owners and custom fields have clear and consistent names. Best use case: creating a project plan and turning decisions into assigned tasks. Example prompt: “Create a plan in Asana for launching a new product page. Divide the work into stages, tasks, dependencies and responsible team members. Show me the proposed structure for approval before saving it.” 11. HubSpot integration with ChatGPT – CRM analysis and record updates The HubSpot integration with ChatGPT provides information about contacts, companies, sales opportunities, tickets and customer interaction history. HubSpot in ChatGPT can analyse the sales funnel, campaign results and customer activity, as well as create and update selected records and log activities. The HubSpot integration with ChatGPT is one of the most extensive solutions available to sales and marketing teams. The quality of its results depends on the completeness of CRM data, consistently defined funnel stages and correctly assigned permissions. Best use case: preparing an account brief, analysing the pipeline, updating an opportunity and creating a follow-up. Example prompt: “Analyse the sales opportunities in HubSpot that have had no activity for 14 days. Identify the priorities and prepare a plan for the next contact.” 12. Salesforce Agentforce Sales integration with ChatGPT – opportunity analysis and CRM management The Salesforce Agentforce Sales integration with ChatGPT combines information about customers, sales opportunities and the pipeline with analysis and planning capabilities. Salesforce in ChatGPT allows sales representatives to prioritise opportunities, prepare account plans, update records and run Agentforce actions directly from a conversation. The Agentforce Sales app for ChatGPT is currently available through the Open Beta programme to eligible customers using the required Agentforce add-ons. Before implementation, organisations should verify their Salesforce edition, access requirements and regional availability. Best use case: preparing a sales representative for a meeting, prioritising opportunities and updating the CRM after a client conversation. Example prompt: “Show me five Salesforce opportunities that require attention this week. Include their value, stage, most recent activity, risk and recommended next step.” 13. GitHub integration with ChatGPT – analysing code, issues and project changes The GitHub integration with ChatGPT gives the model access to context from repositories, code, issues, proposed changes and automated test results. GitHub in ChatGPT can help analyse code changes, organise issues, prepare documentation and identify dependencies between project components. Administrators can specify which repositories the GitHub integration with ChatGPT can access and which operations it can perform. This makes it possible to test the integration on a small number of selected projects before gradually making it available to additional teams. Best use case: analysing proposed code changes, organising issues, reviewing automated test results and preparing change documentation. Example prompt: “Review the open pull requests in the mobile application repository. Identify risks, missing tests and issues blocking the release.” The video shows how to connect GitHub to ChatGPT and give the integration access to selected repositories. The GitHub integration with ChatGPT enables users to explore project structures, analyse code and documentation, and summarise changes, commits and pull requests directly within a conversation. 14. Canva integration with ChatGPT – creating and editing visual content The Canva integration with ChatGPT enables users to search and summarise existing materials, as well as create, edit and display designs directly within a conversation. Canva in ChatGPT is useful for preparing presentations, social media posts, documents and other visual materials. Designs created through the Canva app for ChatGPT remain editable in Canva, allowing the team to continue refining their content and appearance. The best results can be achieved by specifying the intended audience, objective, format, source materials and brand requirements. Best use case: presentations, social media content, sales documents and visual summaries. Example prompt: “Create a presentation in Canva for the management team based on this report. Use eight slides, concise conclusions and one chart on the results slide.” 15. Adobe integration with ChatGPT – editing photos, videos, graphics and PDF documents Adobe for ChatGPT is a package that brings together features from Adobe applications, including Photoshop, Premiere, Firefly, Express and Acrobat. It supports photo editing, consistent batch processing, preparation of social media formats, video shortening, work with PDF documents and searches across Creative Cloud assets. The solution supports workflows intended to produce a finished file. For example, a workflow may begin with a set of employee photos, include lighting correction and consistent cropping, and finish with the export of materials ready for publication. Best use case: repeatable photo editing, adapting content for different channels, working with PDF documents and quickly creating designs from templates. Example prompt: “Standardise the lighting and colours in these photos, apply consistent cropping and prepare versions for employee profiles on the company website.” 15 ChatGPT integrations with business applications – comparison table No. ChatGPT integration Area Best use case Primary type of work 1 Google Drive Documents and knowledge Analysing files from Drive, Docs, Sheets and Slides Search, reading and analysis 2 Microsoft SharePoint Organisational knowledge Working with controlled Microsoft 365 resources Search, reading and analysis 3 Box Content management Secure work with company files and folders On-demand access or synchronisation 4 Gmail Email Summarising email conversations and preparing replies Search, analysis and drafting 5 Outlook Email Microsoft 365 email Analysing email in a business environment Search, analysis and drafting 6 Slack Team communication Finding decisions and summarising channels and messages Search, reading and actions 7 Microsoft Teams Collaboration Analysing conversations, meetings and team context Search and summarisation 8 Notion Knowledge and documentation Creating and updating pages, databases and plans Real-time reading and writing 9 Atlassian Rovo Projects and IT Working with Jira, Confluence, Jira Service Management and Bitbucket Search, creation and updates 10 Asana Work management Managing project portfolios and creating tasks Analysis and project actions 11 HubSpot CRM, marketing and sales Analysing customers, the sales funnel and contact history Analysis, record creation and updates 12 Salesforce Agentforce Sales Enterprise sales Prioritising opportunities, planning accounts and updating the CRM Analysis and sales actions 13 GitHub Software development Working with repositories, issues, proposed changes and automated tests Search, analysis and issue organisation 14 Canva Design and communication Creating editable presentations and marketing materials Search, generation and editing 15 Adobe Creative work and documents Photos, videos, social media content, PDF files and Creative Cloud assets Search, generation, editing and export Security of ChatGPT integrations in a business environment Secure integration of ChatGPT with company systems requires appropriate permission management, separation of read and write operations, selection of the right data access method, approval of actions and operation logging. Data processing terms, OAuth scopes, retention and data residency requirements should also be reviewed for every connected service. In ChatGPT Business, Enterprise and Edu plans, data retrieved through integrations is not used to train OpenAI models. When is it worth building a custom ChatGPT integration? Ready-made ChatGPT integrations cover popular business applications and common use cases. A custom integration becomes justified when critical data is stored in an internal system, the process requires specific logic or the organisation needs greater control over its architecture and information flows. The most common reasons include: a private API, legacy system, internal database or on-premises solution; a workflow involving several systems and rules specific to the organisation; requirements concerning data residency, auditability and approval of operations; the need to combine RAG-based search, business logic and actions performed in external systems; a regulated environment requiring risk assessment, documentation and controlled implementation; a scale at which a custom integration simplifies access and cost management. Such a solution may use a dedicated integration, MCP server, GPT Actions, API layer or an architecture combining several approaches. The starting point should be a specific process, a clearly identified data owner and the expected business outcome. ChatGPT integrations as part of a secure enterprise AI ecosystem ChatGPT integrations provide the greatest value when the connection supports a real process, respects user roles and produces an output that is ready to use. For one organisation, this may mean faster knowledge retrieval. For another, it may involve CRM updates, document automation or a controlled process spanning several systems. Transition Technologies MS designs and implements AI solutions for business tailored to an organisation’s data, architecture, security requirements and operating model. The scope of a project may include API and MCP integrations, RAG solutions, action automation and a model for managing access, risk and accountability. Our approach to AI has been confirmed by ISO/IEC 42001 certification for our Artificial Intelligence Management System (AIMS). TTMS was the first company in Poland to obtain accredited certification for compliance with this standard and is among the first organisations in Europe operating within its framework. This means that we deliver AI projects according to structured principles covering security, accountability, documentation and risk management. TTMS also develops proprietary AI products that support specific business processes: AI4Content analyses documents and creates structured reports; AI4Knowledge helps employees use company knowledge more effectively; AI4E-learning transforms source materials into editable online training courses; AI4Localisation supports the translation and adaptation of content for different markets; AI4Legal automates document analysis and selected legal processes; AML Track supports customer verification, risk monitoring and compliance with AML obligations; AI4Hire structures application analysis and supports the initial assessment of candidates; QATANA uses AI to create test cases and manage the software testing process. This expertise allows us to combine integration, product and regulatory experience. We can help organisations establish a single connection to a company data source or design a solution spanning several systems, access controls and end-to-end process automation. FAQ: Frequently asked questions about ChatGPT integrations Can ChatGPT use multiple connected apps in a single task? Yes. Supported ChatGPT environments can use several approved sources within a single task. For example, a workflow could collect project decisions from Slack, retrieve a report from Google Drive and prepare an action plan in Asana. Availability depends on the ChatGPT plan, the interface or mode being used and the workspace configuration. The prompt should clearly identify the required sources, expected result and point at which ChatGPT should request approval. The organisation should also define which types of data may be combined in a single output. Does an integration give ChatGPT access to all of a user’s data? The scope of access depends on the permissions granted to the integration and the user’s role in the source system. Many integrations respect existing permissions for folders, channels, repositories and CRM records. An administrator account may therefore expose significantly more data than an employee account assigned to a specific team. During configuration, review the OAuth scopes, user roles and options for restricting access to selected resources. A pilot should ideally use an account with permissions corresponding to the intended user role. Can ChatGPT send messages and modify data in external applications? Selected apps, integrations and MCP servers support write actions such as sending messages, creating tasks, updating CRM records or adding pages. The available actions vary by provider, subscription plan and integration version. Some tools show the proposed change and request confirmation before completing it. Administrators may also restrict an integration to read-only access or allow only selected operations. Actions affecting customers, financial data, publications or regulated processes should always have clearly defined human approval requirements. Do I need a paid ChatGPT plan to use integrations? Not always. A limited selection of apps may also be available on the free ChatGPT plan, although search and analysis features may have lower usage limits. Broader access, including data synchronisation and custom MCP-based integrations, usually requires a paid plan such as Plus, Pro, Business, Enterprise or Edu. Availability may also depend on the user’s region, administrator settings and subscription to the external service. The current requirements for a specific integration should be checked directly in the ChatGPT app or plugin directory. Can ChatGPT be connected to a company’s internal system? Yes. An organisation can build a custom MCP server, dedicated integration or API connection that gives ChatGPT access to selected data and actions. A private system may remain behind a firewall or operate on-premises if the architecture uses a secure tunnel and controlled authentication. The project should define tool schemas, roles, logging, action approvals, error handling and protection against prompt injection. Before production deployment, the integration should be tested using valid requests, edge cases and tasks that it is expected to refuse. How do you connect an app to ChatGPT? First, define the task the integration should perform and the data required to complete it. Then open the app or plugin directory in ChatGPT, select the appropriate service and start the configuration process. Sign in to the external application, carefully review the requested permissions and approve only the access that is necessary. Once configuration is complete, open a new conversation, select the connected app and test it using a limited dataset. In a business environment, it is best to begin with a pilot for a small group of users before making the integration available to the wider organisation. Can a company administrator restrict access to apps in ChatGPT? Yes. A workspace administrator can decide which apps and plugins are available within the organisation, who may use them and which actions they can perform. For example, the administrator may allow read-only access while blocking message sending or CRM record updates. In managed workspaces, access can also be assigned according to user roles and groups. Integrations continue to respect permissions in the source system, so users should not gain access through ChatGPT to information they cannot view in the connected application. Does ChatGPT store copies of data retrieved from connected systems? It depends on how the integration works. With on-demand access, data is retrieved when a specific request is processed and is not indexed in advance. Integrations that use synchronisation may create an indexed copy of selected content to speed up searches and improve response quality. Disconnecting an app prevents further access, while its synchronised index is scheduled for deletion from OpenAI systems, typically within 30 days. Information previously used in conversations may remain in chat history, so removing it may also require deleting the relevant conversations and saved memories.

Read
Top 7 AI QA Tools for Pharma in 2026

Top 7 AI QA Tools for Pharma in 2026

In the pharmaceutical industry, test results become part of quality documentation and must be reproducible during an audit. A complete record includes links to requirements, execution details, change history, approvals and audit evidence. When selecting an AI-powered quality assurance tool for pharma, organisations should therefore consider test automation, traceability, data integrity and compliance with GxP requirements. The ranking opens with QATANA, a platform designed for comprehensive test process management. It combines AI capabilities, manual and automated testing, role-based access, audit logs and on-premise deployment. These features address the key needs of pharmaceutical QA teams by accelerating testing, maintaining control over data and supporting complete test documentation. The ranking covers seven solutions addressing different layers of the quality assurance process. Their capabilities include test management and traceability, end-to-end test execution, digital validation, visual testing and device labs. This comparison will help you select a tool suited to a specific system and validation model. Top 7 AI QA Tools for Pharma – Comparison at a Glance Rank Tool Main category Deployment model Best use case in pharma 1 QATANA AI-assisted test management On-premise Controlled testing lifecycle, auditability, and manual and Playwright tests managed in one environment 2 Tricentis Tosca + qTest + Vera Enterprise automation and digital validation Cloud, on-premise or hybrid, depending on the component Large CSV programmes, formal approvals and complex application environments 3 Opkey Enterprise application automation and continuous validation Cloud or on-premise Veeva, TrackWise, Oracle, SAP, Workday and other frequently updated systems 4 Leapwork No-code test automation and continuous validation Cloud, on-premise or hybrid Regression testing of business processes across web, desktop, Salesforce, SAP and Oracle systems 5 Applitools Visual AI and regulated content control Public cloud, private cloud or on-premise Product websites, portals, applications, eIFUs, PDF documents and mandatory safety communications 6 ACCELQ Full-stack no-code automation Public cloud, private cloud, on-premise or hybrid Omnichannel processes covering web, mobile, API, desktop and enterprise applications 7 TestGrid CoTester Agentic testing and device infrastructure Cloud, private cloud or on-premise device lab Mobile applications, patient portals and testing on real devices and browsers What Makes an AI QA Tool Ready for Pharmaceutical Applications? In a regulated environment, test case generation speed is one of several important selection criteria. The tool should support a controlled process in which requirements, risks, test cases, executions, defects and approvals form a consistent chain. Data integrity, decision traceability and the retention of evidence for the required period are equally important. Compliance with GxP, EU GMP Annex 11 and 21 CFR Part 11 depends on how the system is used, configured and maintained within a specific organisation. The assessment should cover procedures, roles, data, supplier qualification and risk analysis. AI capabilities can support this process, while computerised system validation confirms that the solution is fit for its intended use. We assessed AI-powered QA tools for pharmaceutical applications across six areas: Fit for pharmaceutical environments: capabilities, documentation and use cases in pharma, biotech, healthcare or life sciences. Traceability and audit evidence: links between requirements, tests and results, change history, roles, approvals, reports and data exports. Control over AI: review of AI-generated content, execution predictability, change management and human involvement in decision-making. Security and deployment: on-premises deployment, private cloud, data residency, communication with the AI model and access control. Technology coverage: manual, web, mobile, API, desktop, ERP, CRM and legacy application testing, as well as document and device testing. Operational scalability: integrations with Jira, CI/CD and automation frameworks, reporting, licensing models and the effort required to maintain tests. 1. QATANA QATANA combines AI capabilities with features that are particularly important in regulated QA environments. The platform centralises test cases, executions, defects, reporting, and the results of manual and automated tests. This enables teams to manage the testing process and documentation within a single controlled environment. AI generates draft test cases from tickets and requirements and helps select regression suites based on the scope of a given release. In a pharmaceutical environment, generated proposals should be reviewed by the people responsible for requirements, quality and risk assessment. QATANA supports this operating model because every AI-generated item remains an editable test artefact, while the team retains responsibility for its final assessment and approval. QATANA is particularly well suited to pharmaceutical companies that need a central test management system, want to keep data within their own environment and combine manual testing with Playwright automation. During a proof of concept, organisations should verify specific requirements for electronic signatures, retention, artefact versioning and export formats defined in their applicable SOPs. QATANA: for pharma: key facts Tool provider: Transition Technologies MS (TTMS) Website: ttms.com/ai-software-test-management-tool/ Solution type: AI-assisted test lifecycle management platform Key AI capabilities: Draft test case generation, intelligent regression selection, and analysis of ticket data and release information Best use case in pharma: Controlled test management for GxP and non-GxP applications, patient and HCP portals, internal systems and successive software releases Deployment model: On-premises, with the option to configure integration with the organisation’s selected AI model Integrations: Jira, Playwright, AI models and ticketing systems, as well as test artefact import and export Pricing: Custom pricing with a scalable multi-user licensing model What to verify before selection: Signatures and approvals required by SOPs, retention policies, versioning, evidence package exports and AI model governance rules 2. Tricentis Tosca, qTest and Vera Tricentis combines three complementary solutions: Tosca for test automation, qTest for test management and Vera for digital validation and process approvals. The platform supports more than 160 technologies and enables end-to-end testing of processes spanning ERP and CRM systems, web applications, APIs and data layers. The integration of Tosca, qTest and Vera supports requirements management, electronic signatures, formal approvals and the collection of evidence required for CSV. The solution can operate in the cloud, on-premises or in a hybrid model, with the latest agentic capabilities developed primarily for cloud environments. Tricentis for pharma: key facts Tool provider: Tricentis Website: www.tricentis.com Solution type: Ecosystem for enterprise automation, test management and digital validation Key AI capabilities: Agentic test creation from natural language, Tosca Copilot, portfolio and results analysis, and model-based test automation Best use case in pharma: Large CSV programmes, complex end-to-end processes, formal approvals and organisations using multiple enterprise applications Deployment model: Cloud, on-premises or hybrid, depending on the product and required AI capability Integrations: Tosca, qTest and Vera within one process, as well as popular enterprise applications, CI/CD pipelines, APIs, user interfaces and data layers Pricing: Custom pricing based on the selected products, number of users and execution scale What to verify before selection: Required licence scope, availability of AI capabilities in the selected deployment model, data flows and completeness of the validation package 3. Opkey Opkey automates testing for enterprise applications commonly used in life sciences, including Veeva Vault, TrackWise, Oracle, SAP, Salesforce and ServiceNow. Its AI engine analyses the impact of updates, generates test scenarios and automatically repairs tests following interface changes. The platform supports processes spanning multiple systems and provides pre-built libraries of business processes. For GxP applications, it offers IQ, OQ and PQ protocols, electronic signatures, traceability and automated collection of validation evidence. It is particularly well suited to organisations automating the validation of changes in widely used business applications. Opkey for pharma: key facts Tool provider: Opkey Website: www.opkey.com Solution type: No-code test automation and continuous validation for enterprise applications Key AI capabilities: Change impact analysis, test generation, self-healing, root cause analysis and intelligent regression scope selection Best use case in pharma: Validation of updates to Veeva, TrackWise, Oracle, SAP, Workday and Salesforce, as well as processes spanning multiple applications Deployment model: Cloud or on-premises, adapted to the customer’s infrastructure Integrations: Veeva, TrackWise, Oracle, SAP, Workday, Salesforce, Jira, Azure DevOps, qTest, Jenkins, ServiceNow and GitHub Pricing: Custom pricing; demo and test coverage assessment available What to verify before selection: Alignment of pre-built tests with the system configuration, validation protocol content, control over self-healing and maintenance costs following updates 4. Leapwork Leapwork enables visual, no-code automation for web, desktop, ERP, CRM and legacy applications. Its AI capabilities support test generation from natural language, requirements analysis and self-healing while maintaining deterministic scenario execution. The platform’s suitability for GxP environments is demonstrated by its implementation at NecstGen, where 110 workflows were automated and 270 functions within laboratory and quality systems were covered by a compliant process. Leapwork supports cloud, on-premises and hybrid deployment. When selecting a deployment model, organisations should verify the availability of AI capabilities, data processing location and the method used to approve changes proposed by the model. Leapwork for pharma: key facts Tool provider: Leapwork Website: www.leapwork.com Solution type: No-code test automation and continuous validation platform Key AI capabilities: Natural language test creation, self-healing, knowledge building from requirements and documentation, and coverage generation with traceability to source materials Best use case in pharma: Regression automation for web and desktop systems, Salesforce, SAP, Oracle and applications used by quality and operational teams Deployment model: Cloud, on-premises or hybrid Integrations: Playwright, Selenium, Cucumber, GitHub, CI/CD pipelines, test management systems, SAP, Oracle, Salesforce and Microsoft technologies Pricing: Annual subscription with custom pricing based on architecture and execution scale What to verify before selection: Availability status of AI capabilities, data processing location, human approval mechanisms and the ability to freeze a validated configuration 5. Applitools Applitools uses Visual AI to detect visual defects that conventional functional tests may overlook. It compares websites, application screens and PDF documents against approved baselines, identifying issues such as obscured messages, insufficient contrast and incorrect content placement. In pharma, it helps control risk information, instructions for use, regulatory messages and approved product content across devices, markets and language versions. Version history, screenshots, detected differences and approvals create an evidence set that supports QA and compliance teams. Applitools is particularly effective as a visual validation layer supporting functional testing and CSV processes. Applitools for pharma: key facts Tool provider: Applitools Website: www.applitools.com Solution type: Visual AI, visual, functional and cross-browser testing Key AI capabilities: Deterministic visual comparison, detection of significant changes, difference grouping, visual element-based self-healing and root cause analysis Best use case in pharma: Control of approved content, warnings, eIFUs, PDFs, product portals, patient applications and digital accessibility Deployment model: Public cloud, private cloud or on-premises Integrations: Playwright, Cypress, Selenium, Appium, more than 50 frameworks, Jira and popular CI/CD tools Pricing: Free trial; Starter and Enterprise plans priced individually What to verify before selection: Baseline approval rules, retention of screenshots and detected differences, language version support, audit package exports and the scope of accessibility testing 6. ACCELQ ACCELQ is a no-code platform for testing web, mobile, API and desktop applications, as well as enterprise systems. Its AI capabilities support scenario design, change impact analysis, self-healing and automation maintenance. The platform can test processes spanning multiple systems, such as portals, APIs, Salesforce, SAP and Oracle. SaaS, private cloud, on-premises and hybrid deployment models allow organisations to align the architecture with their data processing policies. ACCELQ for pharma: key facts Tool provider: ACCELQ Website: www.accelq.com Solution type: Unified no-code platform for test management and full-stack automation Key AI capabilities: Scenario generation, process modelling, change impact analysis, self-healing and AI-assisted automation maintenance Best use case in pharma: End-to-end processes spanning web, mobile, API, desktop, backend, Salesforce, SAP, Oracle and other enterprise applications Deployment model: Public cloud, private cloud, on-premises or hybrid Integrations: Jira, Azure DevOps, Jenkins, GitHub, GitLab, TeamCity, Bamboo, Salesforce, SAP, Oracle and Workday Pricing: Annual subscription with custom pricing; a 14-day free trial is available What to verify before selection: Validation documentation package, signatures and approvals, complete AI data flow, and the cost of private cloud or on-premises deployment 7. TestGrid CoTester TestGrid combines the CoTester agent with a cloud of real devices and browsers and a private device lab. AI generates tests from requirements or an application URL, updates them following interface changes and allows users to approve each scenario before execution. The platform supports web, mobile, API, visual and performance testing, as well as existing Selenium, Appium, Cypress and Playwright test suites. In pharma, it can support the testing of patient portals, therapeutic applications and solutions used in clinical trials on real devices. For on-premises deployments, organisations should note that the AI capabilities require a connection to TestGrid’s hosted infrastructure. TestGrid CoTester for pharma: key facts Tool provider: TestGrid Website: www.testgrid.io Solution type: Agentic testing, test management, and cloud or on-premises device lab Key AI capabilities: Test generation from requirements, conversational editing, AgentRx self-healing, error summarisation and results analysis Best use case in pharma: Mobile and web applications, patient portals, field solutions and testing on real devices and browsers Deployment model: Cloud, private cloud or on-premises device lab; AI capabilities may require an outbound connection Integrations: Jira, Jenkins, GitHub Actions, GitLab, Azure DevOps, Selenium, Appium, Cypress and Playwright Pricing: Starter plan from USD 199 per user per month, based on pricing available in August 2026; Growth and on-premises plans are priced individually What to verify before selection: Scope of data sent to the hosted AI service, data residency, log immutability, retention policies and the ability to operate without external connectivity AI-generated image. The people depicted are fictional. Before Implementing an AI QA Tool in Pharma: 9 Questions to Ask the Vendor Before selecting a tool, conduct a proof of concept under conditions that closely reflect the actual testing process. This allows you to assess test creation and execution speed, documentation completeness and the ability to reconstruct the entire process during an audit. Before selecting and implementing an AI QA tool in a pharmaceutical company, ask the vendor the following questions: Does the platform connect requirements, test cases, test executions and reported defects? Which activities and changes are recorded in the audit logs? Can roles and permissions be configured according to the organisation’s procedures? Can AI-generated test cases be reviewed, edited and approved before use? Are manual and automated test results available in one consistent view? How does on-premises deployment work, and how can the platform connect to an AI model selected by the organisation? Does the platform integrate with the organisation’s existing tools, such as Jira, Playwright and CI/CD pipelines? Can test data, reports and other artefacts be imported and exported in the required formats and scope? How do licensing, deployment, integrations, training and ongoing support affect the total cost of the solution? Best AI QA Tool for Pharma: Final Recommendation The final decision should reflect the intended use, risk assessment and a proof of concept conducted on a representative process. The solution provider also plays an important role, as its experience affects implementation quality, change management and the audit readiness of the testing process. TTMS, the provider of QATANA, has worked in the pharmaceutical industry since 2011, involving more than 400 specialists in over 100 projects and services. The company combines QA engineering with expertise in quality management and computerised system validation in line with GAMP 5 and EU GMP Annex 11. These capabilities are supported by the TTMS Integrated Management System, which includes ISO 9001 and ISO 27001. This enables TTMS to support the entire implementation lifecycle, from requirements definition and tool configuration to validation, maintenance and controlled change management. FAQ Is a “21 CFR Part 11 compliant” claim sufficient when selecting an AI QA tool for pharma? No. Such a claim usually describes the available features or the way the product has been designed, while compliance is assessed for a specific intended use and implementation. The organisation must determine which electronic records and signatures fall within scope, configure roles, permissions, audit trails, retention policies and procedures, and then demonstrate that the system is fit for its intended use. Integrations with Jira, CI/CD pipelines, code repositories and other systems are also important because data flows may extend beyond the QA tool itself. Vendor documentation can facilitate validation, but responsibility remains with the pharmaceutical company. A proof of concept should therefore include the reconstruction of a complete evidence chain from the original requirement to the approved test result. How should AI-generated test cases be validated in pharma? An AI-generated test case should be treated as a draft requiring expert review. A person familiar with the requirement and its associated risk should verify the preconditions, test data, steps, expected results, negative scenarios and traceability to the source requirement. The system should record the source, model version, generation date, approver and all subsequent changes. Functionality with a greater potential impact on product quality, patient safety or data integrity requires more rigorous review and independent approval. AI performance should be evaluated against a controlled reference set using measures such as coverage completeness, the number of rejected suggestions and errors identified during review. This approach preserves the time-saving benefits of AI while keeping accountability with qualified personnel. Are self-healing tests safe in a validated GxP environment? They can be used when the mechanism operates in a controlled manner and maintains a complete record of every change. Automatically correcting a technical locator can reduce false failures, provided that the repair does not alter the meaning of a step, the acceptance criterion or the scope of the test. A well-configured system displays the proposed change, its rationale, and the previous and new values, and requires approval for significant modifications. The organisation should define in its SOPs which repairs may be accepted automatically, which require review and when a test must be reapproved. False positives, false negatives and the effects of self-healing engine updates should also be reviewed periodically. Execution repeatability and decision traceability are more important than the number of tests repaired without tester involvement. Can production data from pharmaceutical systems be sent to an external AI model? The preferred starting point is to use synthetic, anonymised or masked data limited to the minimum required for testing. Sending production data requires a legal basis, information classification assessment, vendor agreement, transfer controls, retention rules, processing location controls and clear policies regarding model training. Patient data, clinical trial information, safety data and confidential product information require particular protection. An on-premises deployment may still rely on a hosted AI service if an agent or interface communicates with an external model. The architecture should therefore show separately where tests are stored, where automation is executed and where AI processing takes place. Access to sensitive data should be granted only after the vendor provides a clear and verifiable description of the complete data flow. What should be done after an update to the AI model used by a QA tool? A model update should be managed as a controlled change, with the scope of assessment determined by risk. The first step is to identify which functions use the model and whether the update could affect test generation, regression selection, self-healing, defect classification or reporting. A previously approved reference test set should then be executed, and the results compared with those produced by the earlier version. Any differences should be assessed, documented and approved before the updated model is used more broadly. The change record should include the model version, date, scope, assessment results, accepted limitations and the person responsible for the decision. When a vendor updates the model without offering the option to freeze a version, the agreement should define advance notification, a testing window and a rollback procedure. Effective model version control is essential for maintaining the validated state.

Read
Astra, the Future GPT-6? OpenAI’s New Model Explained

Astra, the Future GPT-6? OpenAI’s New Model Explained

Solving mathematical problems that scientists had wrestled with for years – could there be a better demonstration of what a new AI model can do? OpenAI has typically previewed new versions of its large language models with benchmark results, meaning scores from standardised tests designed to measure a model’s capabilities. I have to admit that seeing GPT tackle genuine research problems makes a much stronger impression on me. What will you learn about OpenAI Astra? What Astra is and why it is being discussed as a potential GPT-6, 10 results in mathematics and theoretical computer science presented by OpenAI, How Astra analyses problems, tests hypotheses and changes its approach, The differences between a conversational model, an AI agent and a system capable of managing an entire project, What Astra could mean for science, business and the future of AI models, Critical responses to the model’s achievements, Cybersecurity risks associated with autonomous AI agents, Which important questions OpenAI has yet to answer. What is OpenAI Astra, and could it become GPT-6? OpenAI describes Astra, the prototype’s working name, as “our next major model”, although the company has disclosed very few details so far. Its task was to develop arguments independently, test hypotheses, recognise unproductive approaches and find new paths towards a solution. The results of its work can then undergo formal and independent verification. We do not know how Astra is built, how much information it can analyse at once or how it organises its work on a complex task, although we can speculate about the last of these. OpenAI has also not disclosed whether Astra is a single model, a team of collaborating AI agents or a more extensive system equipped with mechanisms for coordinating their work and retaining previous results. The prototype may be connected to a model previously described by OpenAI as capable of operating autonomously over very long periods. Such a system can make repeated attempts, analyse intermediate results and maintain its direction of work for many hours, potentially even days. According to media reports, Sam Altman has already presented Astra to US politicians and regulators. The term “GPT-6 Astra” should therefore be treated as media shorthand. Astra could eventually be released as GPT-6, another version of GPT-5 or a separate family of models. For now, all of these possibilities remain open. Why could Astra’s 10 results matter more than another benchmark record? OpenAI presented ten results concerning problems that had remained open for at least a decade and, in most cases, considerably longer. The problems come from eight fields: high-dimensional geometry, coding theory, group theory, operator algebras, computational complexity theory, quantum computing, lattice geometry and post-quantum cryptography, extremal combinatorics. In simple terms, the process worked as follows: GPT generated mathematical arguments. Once the results had been obtained, researchers worked with the model to develop them into scientific papers. The system then translated the arguments into Lean 4, allowing a computer to check every step of the proofs. For readers interested in the technical details, here are the relevant links: the complete collection of papers, the Lean formalisation repository and reconstructions of how the solutions were developed. Independent verification of all the claims by the scientific community is only beginning. Mathematicians can now review the papers, check the definitions, run the formalised proofs and look for potential gaps. I discuss this in more detail in one of the final sections. 10 new results from Astra in mathematics and theoretical computer science A quick warning: this section is about to become fairly technical. These subjects are new, abstract and extraordinarily difficult for me as well, so I have tried to explain each result in the simplest possible terms. Here is how GPT Astra approached the individual problems. 1. Sphere packing in high-dimensional spaces The sphere-packing problem asks how densely identical spheres can be arranged, much like coins on a table or balls in a box. Mathematicians also study this question in spaces with hundreds or thousands of dimensions because it has applications in areas such as information theory and data encoding. Astra used an established mathematical method to determine more precisely how densely spheres can be packed in spaces with a very large number of dimensions. According to the authors, this is the first improvement since 1978 to the value used in the formula describing how quickly the possible packing density decreases as the number of dimensions increases. The difference becomes more significant as the number of dimensions grows and enables a more precise estimate of the maximum packing density. Put simply, Astra’s calculations improve our understanding of how many spheres can fit inside such a “high-dimensional box”. 2. Binary and spherical codes: new bounds on the number of error-resistant codes A binary code is a set of sequences made up of zeros and ones. These sequences must differ from one another sufficiently for a system to detect and correct transmission errors. This can be compared to positioning transmitters at safe distances from one another so that their signals remain easy to distinguish. Astra determined more precisely how many codes can be placed sufficiently far apart for a system to continue distinguishing between them and correcting errors. This enables mathematicians to estimate more accurately how many codes with the required level of error resistance can fit within a given space. The model tested its initial idea on a simple example consisting of eight digits and discovered that it produced an incorrect result. It therefore abandoned that approach and reformulated the problem. This case demonstrates Astra’s ability to test its own assumptions and redesign its solution when the original direction proves unsuccessful. OpenAI’s published materials support four important conclusions. 3. The first explicit example of a non-sofic group A group is a mathematical way of describing symmetries and operations that can be performed in sequence, much like a set of moves used to rotate a Rubik’s Cube. Sofic groups can be approximated with arbitrary precision using simpler structures based on a finite number of elements. For decades, mathematicians wondered whether this property applied to every group. Astra identified a specific example of a group that cannot be approximated with arbitrary precision using simpler models composed of a finite number of elements. The result demonstrates that these simplified models cannot represent every mathematical group. The solution combined several distant areas of mathematics, demonstrating the model’s ability to bring together tools that had not previously formed an obvious path towards a proof. 4. Disproving Connes’ rigidity conjecture The von Neumann algebra associated with a group can be compared to its highly complex mathematical “fingerprint”. Connes’ conjecture proposed that, for a certain class of particularly rigid groups, this fingerprint uniquely identifies the group in question. Astra constructed infinitely many different groups with exactly the same mathematical “fingerprint”. In doing so, it disproved Connes’ conjecture and answered a later question posed by mathematician Sorin Popa. Astra used a mechanism resembling the carrying operation in binary addition. This made it possible to construct many different groups with the same mathematical “fingerprint”. Put simply, Astra demonstrated that a single mathematical “fingerprint” can belong to infinitely many different groups. 5. The matrix permanent: the minimum number of operations required for its computation The permanent of a matrix is calculated in a similar way to the determinant, except that all terms are added with a positive sign. This seemingly minor change makes the permanent one of the most important examples of a problem with extremely high computational complexity. Astra determined the minimum number of basic operations required to calculate the permanent. It proved that no solution within this class can be simplified below a certain level of complexity. This can be compared to determining the minimum number of components required to build any machine capable of performing a particular task. Such a proof must cover every possible construction that meets the specified conditions, which makes it exceptionally difficult to develop. The result provides a more precise lower bound on the number of operations needed to solve this problem. It also brings mathematicians closer to answering a fundamental question: which problems can be solved efficiently, and which will always require an enormous amount of computation? 6. Quantum games: why does the probability of a perfect win decrease so rapidly? Imagine a game in which two players answer a referee’s questions separately, while their shared goal is to complete every round successfully. In the classical version, each additional round rapidly reduces the probability of a perfect win, much like repeatedly tossing a coin reduces the chance of getting heads every time. In the quantum version, the players’ results can be correlated even when they do not communicate during the game. They can also analyse several rounds as a single combined problem. Astra proved that even such quantum correlations cannot prevent the probability of winning every repeated round from decreasing very rapidly. The problem had remained open since at least 2004. The key to the solution was a method for transforming quantum states without changing the probabilities of their possible outcomes. The result advances the theory of interactive proofs, quantum information theory and methods for increasing the reliability of protocols. 7. The Closest Vector Problem: even an approximate solution remains difficult A lattice can be imagined as a regular grid of points, similar to street intersections in a perfectly planned city, extending across many dimensions. The Closest Vector Problem (CVP) involves finding the point on this grid that lies closest to a selected location. It is highly relevant to geometry, coding theory and post-quantum cryptography. Astra connected CVP with the well-known 3SAT logic problem and demonstrated that finding even a solution that merely approximates the optimal one is extremely difficult. This difficulty increases with the number of dimensions in the lattice. The model represented the logical puzzle as a system of points and distances between them. This can be compared to encoding a complex logic puzzle in a spatial arrangement of points so that solving one problem also provides a solution to the other. The result deepens our understanding of the theoretical difficulty of lattice-based mathematical problems. Assessing the security of specific cryptographic algorithms requires a separate analysis of their variants, parameters and methods of data generation. 8. Proving Ehrhart’s conjecture on the volume of high-dimensional shapes A high-dimensional convex body can be imagined as a solid placed on a regular lattice of points, with its centre of gravity being the only lattice point located inside it. Ehrhart’s conjecture specified the maximum possible volume of such a body, and Astra proved it for any number of dimensions: vol(K) ≤ (n+1)n / n! The main difficulty was connecting the number of individual lattice points with the volume of the entire body. The first approach provided only part of the information required. Astra therefore reformulated the problem in the language of another branch of geometry and began searching for a solution using its tools. The model combined several advanced methods for describing the body’s shape, its boundaries and the distribution of points. Put simply, Astra translated the geometric puzzle into a different mathematical language in which it became possible to determine the exact volume bound. 9. Multicolour Ramsey numbers A complete graph can be imagined as a group of people in which every pair is connected by a line, with each line assigned one of k colours. Mathematicians ask how large such a network must become before it inevitably contains three people whose connecting lines are all the same colour. This minimum size is denoted by Rk(3). Astra developed new colouring methods which, when combined with previous results, established the growth rate of this number: Rk(3) = kΘ(k). The result does not provide an exact value for every number of colours, but it reveals the correct scale of growth. In doing so, it resolves Erdős Problem No. 183. Astra expanded the network in stages according to the same rule. This made it possible to construct increasingly large configurations without creating a triangle whose edges were all the same colour. 10. Two counterexamples in extremal graph theory An extremal number determines how many connections a network can contain before a specified forbidden configuration inevitably appears. Astra disproved two conjectures proposed by Erdős and his collaborators concerning how this value could be predicted. In the first case, it constructed a family of graphs in which forbidding each member individually still allowed approximately n4/3 edges, while applying all the restrictions simultaneously reduced the maximum number to O(n21/16). This shows that several forbidden structures can constrain a graph far more strongly together than when each is considered separately. In the second case, Astra found a network divided into two groups in which every small section contained few connections, while the complete construction could be considerably denser than the conjecture predicted: ex(n, H) ≥ cn3/2+ε. The two results resolve Erdős Problems No. 146 and 180. They also show that the simple structure of small sections of a network does not always allow us to predict how dense the entire construction can become. A critical perspective: how do experts assess Astra’s mathematical achievements? After the initial excitement, important reservations began to emerge. Mathematicians pointed out that at least two of Astra’s results rely heavily on earlier work, raising questions about their novelty. OpenAI has since changed the way it describes the experiment. It now increasingly refers to “making meaningful progress”, rather than solely to “solving longstanding problems”. Interestingly, a researcher affiliated with Anthropic reported that the Claude Fable model had reproduced solutions to five of the ten problems tackled by “GPT-6” within 24 hours, although these results have yet to be fully verified. This does not undermine Astra’s capabilities, but it makes it more difficult to determine whether we are witnessing a breakthrough driven by the exceptional abilities of one model or broader progress across AI models as a whole. Above all, there is still no reliable, independent and fair comparison conducted using the same problems, prompts, computational budgets and rules governing access to tools. What do the results reveal about how Astra works? OpenAI’s published materials support three important conclusions. 1. The model can abandon dead ends The published reconstructions show Astra trying different approaches, identifying obstacles, reformulating problems and returning to earlier stages of its work when necessary. This resembles genuine research more closely than an extended answer generated in a single pass. In the binary-codes problem, the first recurrence was rejected after the model found a small counterexample. When working on Ehrhart’s inequality, Astra spent considerable time developing an approach based on symmetrisation before reformulating the problem in terms of toric geometry. In the proof concerning quantum games, it recognised that the classical argument lost control after conditioning on rare events and began searching for a representation that preserved quantum probabilities. The published document does not reveal the model’s complete internal reasoning process. It is a narrative produced by a model that reviewed the original reasoning traces and the final papers. 2. GPT Astra combines discovery with automated verification Lean checks the correctness of a formal proof step by step. The repository contains separate files for all ten results, along with instructions for performing additional checks of the formalised proofs. Computer verification does not replace assessment by independent mathematicians. Researchers must still establish, among other things: whether the formal theorem corresponds precisely to the original problem, whether the definitions introduce any unintended simplifications, whether the result is genuinely new, how significant it is for the relevant field, whether the manuscript correctly connects the formalisation with the informal argument. A computer can confirm that a written proof is logically correct under the adopted definitions and assumptions. It does not automatically confirm that the authors formalised precisely the version of the problem that mathematicians intended to address. The research was published on 1 August 2026, so full independent verification by the mathematical community will take time. Thomas Bloom of the University of Manchester nevertheless described the results as “big news” and rated the significance of the presented constructions particularly highly. 3. The cost of Astra’s results and the importance of additional computing power OpenAI claims that, based on the API pricing for GPT-5.6 Sol, the tokens required to find all ten solutions would have cost approximately $2,000. This figure is, of course, neither the actual cost of developing Astra nor the full cost of the project. It does not include model training, infrastructure, researchers’ work, problem selection, validation or all the unsuccessful attempts. It is simply the cost of the tokens used to find the published solutions, calculated according to current API pricing. The average comes to approximately $200 per published result, but we do not know: the total number of problems presented to the model, the success rate, how the costs were distributed across the problems, how long the system operated, how many agents were involved, how many runs were conducted in parallel. Noam Brown, an OpenAI researcher involved in the work on Astra, acknowledged that the system had also been tested unsuccessfully on other major mathematical challenges, including the Millennium Prize Problems. He added that OpenAI had not allocated an especially large amount of computing power to each problem. The company therefore believes that Astra could achieve better results if given more time and resources to search for solutions. From GPT-5.6 to Astra: how AI is moving from answering questions to managing projects GPT-5.6 already includes several features that point towards the direction described above. The model can independently select tools, analyse the results it obtains and use them to plan its next actions. Ultra mode uses four agents by default, while OpenAI has also tested configurations involving sixteen agents. The company also offers a multi-agent mode in the Responses API in beta. Astra may develop this architecture towards much longer and more coherent periods of autonomous operation. The most important difference would be its ability to manage an entire project over many hours or days. The system would need to remember what it had already tried, which ideas it had rejected, what results it had obtained and how the individual tasks related to one another. From the user’s perspective, the change could be very tangible. Instead of guiding the model through a sequence of prompts, the user gives it an objective, a set of available tools, a defined scope of permissions, a budget and completion criteria. The user then returns to a finished result accompanied by a record of the attempts, tests and decisions made along the way. This progression can be presented as three successive units of work: A conversational model generates an answer. An agent completes a task using tools. A multi-agent system manages a project in which tasks are created and modified as the work progresses. Only the technical documentation will show whether Astra genuinely operates at the third level as a coherent system. The mathematical demonstration is, however, the first strong indication that this direction is becoming more than a promise. OpenAI, Google DeepMind and Anthropic: the race to develop long-horizon AI models Google DeepMind, Anthropic and OpenAI are developing AI systems capable of independently handling increasingly long and complex tasks. Aletheia, Google’s mathematical agent based on Gemini Deep Think, can generate solutions, verify their correctness and revisit them when it detects an error. When analysing 700 Erdős problems, it solved four questions that had previously remained open. Anthropic, meanwhile, is focusing on coordinating the work of multiple agents. According to the company, Claude Opus 4.8 can divide a large project into smaller parts and assign them to hundreds of subagents working in parallel. This allows it to carry out tasks such as migrations involving hundreds of thousands of lines of code. Claude Science, another environment being developed by the company, is intended to make it possible to trace and verify the successive stages of research work. All these projects point in the same direction: models are expected to work towards a single objective for longer, monitor their own results and revise earlier decisions. Astra stands out for producing results at the frontier of contemporary knowledge and for formally encoding some of its proofs, allowing their correctness to be checked by a computer. How could Astra change the AI model and agentic tool market? 1. Benchmarks may lose their role as the primary evidence of AI model quality Competition will increasingly focus on the final outcome: a new hypothesis, a discovered vulnerability, a completed system migration, a developed scientific model, a working application, a result that can be verified automatically. Astra was presented through its scientific results because conventional benchmarks do a poor job of communicating the difference between a model that answers a question and a system that manages an entire project. Benchmarks will remain necessary for comparing models under controlled conditions. Their market significance may, however, decline in favour of evaluations that measure project completeness, operational continuity and the quality of the final result. 2. The cost of a completed task may matter more than the price per token For business customers, the following factors will become increasingly important: the cost of completing the project, the time required to obtain the result, the probability of success, the number of human interventions, the cost of validation, the ability to resume work after an error. Approximately $2,000 for the tokens that produced ten published results is a powerful economic signal, even with all the reservations concerning how the results were selected. Perhaps future pricing models will include offers such as the “cost of a successfully completed migration”. This would, however, require transparent information about the number of failed attempts, the additional human work involved and the cost of verifying the result. 3. Astra could affect platforms that coordinate AI agents If models begin dividing work among agents, retaining a record of their progress and combining the resulting outputs themselves, some external platforms used to coordinate AI agents may lose part of their current advantage. The value of these platforms will primarily come from integration with company systems, access and cost controls, visibility into agent activity and the ability to audit the entire workflow afterwards. Dividing a task among several agents may become a standard model capability. The implementation layer will derive its value from embedding agents securely within organisational processes. 4. Keeping AI agent workflows under full control When a task lasts several days, the system must regularly save its progress and check whether subsequent actions are still leading towards the defined objective. OpenAI acknowledges that long-running work requires monitoring the model’s complete sequence of actions. Each individual step may appear safe even though, taken together, the steps lead to an unacceptable outcome. The company described a case in which a model attempted to bypass a safeguard by splitting a digital access key into separate parts and reassembling them only when the program was launched. Each step, considered in isolation, could appear relatively harmless, while the objective of the complete sequence was unacceptable. In another experiment, the model was instructed to deliver its results through Slack. Instructions found in the code it was analysing suggested publishing the proposed changes directly to GitHub instead. The model spent approximately an hour searching for a way around the test environment’s restrictions before finding a vulnerability and publishing the result there. Earlier models had abandoned the task under similar circumstances. The case shows that a longer operating horizon increases both the model’s effectiveness and the amount of time it can devote to searching for ways around safeguards. In response, OpenAI developed a system that analyses the model’s entire sequence of actions as its work progresses. The system can stop the session, notify the user and ask them to decide whether the work should continue. Further details are available in OpenAI’s report on the safety of long-horizon models. 5. Research fields in which Astra could accelerate progress The most immediate impact is likely to appear in fields with: precisely defined problems, extensive available literature, formal or automated verification tools, the ability to conduct computational experiments, unambiguous criteria for measuring progress. Mathematics is an ideal testing ground because a proof can be verified. Similar conditions exist in software development, chip design, some areas of chemical research, bioinformatics and cybersecurity. Economics, strategy, law, management and social research will remain much more challenging because correctness cannot be reduced to a machine-verifiable certificate. In these fields, a model may produce an impressively coherent project that is still based on flawed assumptions or a poorly defined objective. 6. Long-horizon models will require more computing power An important capability of a model will be the option to allocate more computing power and more attempts to particularly difficult problems. This will give an advantage to laboratories with: extensive computing resources, efficient communication between agents, effective context management, automated detection of dead ends, the ability to run multiple attempts and select the best result. The next stage of competition may concern more than model size. It may also depend on how effectively models use time and computing power when working on a specific task. The same model could operate as a relatively inexpensive assistant for everyday questions and as a costly research system when the user increases the budget for time, agents and parallel attempts. The section likely to age quickly: what do we still not know about Astra? OpenAI has not disclosed basic information about Astra, including its architecture, size, method of agent collaboration, memory mechanism or capabilities beyond mathematics. We also do not know its price, release date or whether OpenAI plans to make the model available through ChatGPT or the API. The published results do not demonstrate that Astra selected the problems independently, operated without supervision or can manage an entire research process. Nor do we know whether it can achieve similar results in other fields. There is therefore no basis for describing Astra as a system that matches human capabilities across a broad range of intellectual tasks. We also do not know the total number of failures. OpenAI published selected successes, while Noam Brown confirmed that the system had attempted to solve other major problems without success. Without knowing the total number of attempts, it is impossible to calculate Astra’s actual success rate or the expected cost of obtaining one valuable result. Why is Astra not yet an autonomous scientist? The published papers show a system solving problems selected and presented by humans. An autonomous scientist would also need to: select research directions independently, assess which questions are important, determine whether a result is genuinely new, design subsequent experiments, decide when sufficient evidence has been collected, place the result within the broader context of the field. Astra completed the most technically demanding part of this process: it developed new arguments and brought them to a form that could be formally verified. This is a major achievement, but it does not encompass the full scope of scientific work. Can Astra succeed beyond mathematics and controlled environments? The ten published papers demonstrate what Astra was able to achieve in a carefully selected environment. Mathematics offers clearly defined problems, extensive literature, precise language and formal verification tools. The real test will be whether this capability can be transferred to projects in which the objective changes as the work progresses, tools fail, data is incomplete and the correctness of the result requires human judgement. If Astra can maintain a coherent process over many hours or days, delegate subtasks, retain the results of previous attempts and return to a problem after detecting an error, the change will be more significant than another increase in benchmark scores. Models such as Astra demonstrate how rapidly the capabilities of artificial intelligence are advancing. In business, their value depends on selecting the right process, ensuring data quality, integrating AI with company systems and maintaining control over its operation. TTMS helps organisations design and implement solutions tailored to specific operational needs. Explore TTMS AI solutions for business and implementation examples. How autonomous was Astra when solving mathematical problems? OpenAI states that the mathematical arguments were generated by the system, while humans contributed to preparing the manuscripts, formalising the results and verifying their correctness. The papers list OpenAI as the author, and the company has not attributed individual proofs to specific employees. This creates an interesting precedent: the organisation assumes responsibility for the publications while crediting the model with producing the arguments. However, it remains unclear who selected the problems, prepared the prompts, initiated subsequent attempts and decided which results were suitable for publication. Without this information, it is difficult to determine Astra’s precise level of autonomy or distinguish the capabilities of the model itself from the work of the wider research team. Can artificial intelligence be the author of a scientific paper? Authorship involves responsibility for the research method, the evidence presented, the conclusions and any potential errors. An AI system cannot formally accept such responsibility, so researchers should remain the authors of scientific publications. The model’s contribution should be described clearly in the methodology, including how it was used and which elements of its work were verified by humans. How can researchers verify whether AI has made a genuinely new discovery? A correct result is not necessarily a new one. Researchers must compare it with the existing literature, previously unpublished work and known variants of the same problem. One particular challenge is determining whether the model developed a new solution or reproduced a relationship contained in its training data. Novelty should therefore be assessed separately from the correctness of the proof itself. Can a result produced by a closed AI model be reproduced? Reproducing an experiment is difficult when researchers do not know the model’s architecture, training data or exact settings. Recording the prompts, system version, tools used, intermediate results and human interventions can make the process more transparent. The final result should also be verifiable using a method independent of the model that generated it. Without this documentation, other scientists may be able to verify the result itself, but not the full process that led to it. Could AI agents increase the risk of errors and unreliable scientific publications? An AI agent can generate large numbers of convincing hypotheses, proofs and interpretations of data in a short time. This scale can accelerate research, but it can also spread flawed assumptions more quickly. Academic journals and research institutions will need clear rules for disclosing the use of AI, preserving a record of the research process and independently verifying the most important results. The transparency of the process will become as important as the quality of the final publication. How should a research team prepare to work with AI agents? A good starting point is to select tasks with results that can be verified unambiguously. The team should determine which data and tools the agent can access, which actions require human approval and who is responsible for accepting the final result. It should also establish procedures for recording each stage of the work, reporting errors and stopping an experiment when necessary. This preparation allows researchers to benefit from the speed of AI while maintaining control over the quality of the research.

Read
GPT-Powered AI Agents: How to Match Autonomy to the Process?

GPT-Powered AI Agents: How to Match Autonomy to the Process?

Until recently, enterprise automation followed a simple division: systems performed tasks defined by rules, while cases requiring interpretation were passed to people. GPT-powered AI agents expand the range of processes that can be supported through automation. They can work with documents, incomplete data and the language used by customers or employees, making them suitable for processes that were previously difficult to automate. For large organisations, this raises a practical question about AI agent autonomy: where does expert support end, and where does independent action within a process begin? In some situations, the agent’s role is to gather information and prepare a recommendation. In others, it prepares an action for approval. There are also areas where it can independently carry out repetitive steps when the organisation has defined the rules, permissions, limits and exception-handling paths. GPT-powered AI agents can already support teams with ticket handling, document analysis, decision preparation, data updates and multi-step tasks. The key implementation question is: which decisions and actions should remain with people, and which can an agent perform within agreed rules? An AI agent in the enterprise is a process participant, not just a chatbot In practice, a GPT-powered agent needs five elements: access to reliable sources of knowledge, a clearly defined business objective, tools and integrations with enterprise systems, permissions aligned with its role, rules that define the boundaries of its actions. A language model can interpret the content of a document, a customer message or an incident description effectively. It does not, however, replace a business process. Workflows, permissions, validations and decision history are what make an agent operate predictably, even when it handles hundreds or thousands of cases each month. Three levels of AI agent autonomy In a large organisation, it is worth designing agents across three levels. This allows autonomy to grow alongside process maturity and trust in the solution. Operating level Agent’s role Example tasks Human role Level 1: Advisory agent Analyses information and prepares a recommendation. Case summary, risk identification, proposed response, ticket prioritisation. Makes the decision and carries out the action. Level 2: Agent preparing an action for approval Completes the next steps in a process, stopping before actions with significant consequences. Creates an application, updates data, prepares a communication, submits an instruction for approval. Reviews and approves specified steps. Level 3: Agent performing tasks automatically Independently carries out tasks in line with the process policy. Case classification, status updates, sending standard information, creating a task in a system. Handles exceptions, monitors quality and updates process rules. The level of autonomy does not need to apply to the entire agent. The same agent may independently classify tickets, prepare a response that requires approval and transfer unusual cases to an expert. In practice, an organisation therefore designs autonomy for individual decisions and actions, rather than choosing a single operating model for the whole solution. What determines whether an AI agent can complete a task independently? A useful starting point is to assess two factors: the impact of the action on the organisation and whether it can be reversed. The greater the business, legal, financial or reputational consequences of a decision, the more important human approval becomes. Nature of the action Recommended model Low impact, simple rules, easy to reverse Automatic execution with a record in the process history. Medium impact, data from several sources, possible exceptions The agent prepares the action and an authorised person approves it. High financial, legal or customer impact The agent presents analysis, options and justification. The decision remains with a person. Unclear rules, incomplete data or conflicting information Automatic escalation to an expert, together with the context and collected data. This principle is particularly useful in organisations operating across multiple countries, with complex permission structures and a large number of systems. Just as important as the list of tasks is knowing what the agent must not do and when it should hand a case over to a person. 7 questions to ask before giving an AI agent permission to act What action should the agent perform? Describe it specifically, for example: “create a service ticket”, “update contact details” or “prepare a response to a complaint”. What data will it work with? Identify the sources, data owners, update frequency and access rules. What business rules must it follow? These may include financial limits, contractual terms, SLA levels, compliance requirements or communication policies. What exceptions should stop the process? The agent needs a clear escalation path for unusual or incomplete cases, or those requiring specialist assessment. Can the action be reversed? The ease of correction affects the appropriate level of autonomy, the scope of testing and the need for additional approval. Who is accountable for the decision? The process owner, approver and technical team should all have clearly assigned roles. How will the organisation establish why the agent took a particular action? The case history should show the input data, rules, sources used, recommendation and process outcome. This is why AI agent projects often begin with bringing the process itself into order. The organisation gains more than a new AI capability: it also gains better visibility of responsibilities, exceptions and how work actually flows. Where can GPT-powered AI agents add value in a large enterprise? Customer service and back-office teams An agent can read a customer message, identify its subject, retrieve data from a CRM or case-management system, prepare a response in line with company policy and route it to the appropriate queue. For standard cases, it can also update a status, create a task for the team or send the customer a confirmation. Full autonomy works well for low-risk actions, such as providing information about the status of a ticket. Complaints, individual commercial terms or cases requiring interpretation of a contract should be passed to an employee together with the agent’s analysis. Finance, procurement and document workflows An AI agent can read a document, check whether the data is complete, compare it with a purchase order and flag discrepancies that require clarification. It can also prepare a case summary, collect missing information and initiate the appropriate approval workflow. Decision thresholds are particularly important in this area. The agent can process a document automatically when it meets all conditions, while cases that exceed a defined amount, contain discrepancies or concern a new supplier can be submitted for approval. IT, administration and ticket management In an IT environment, an agent can classify tickets, create an incident summary, search for similar cases in the knowledge base, propose actions in line with a runbook and update the user on progress. In administrative processes, it can prepare an application, complete data in a form and remind the requester about missing documents. For actions involving configuration changes, access permissions or production systems, an approval-based model is advisable. The agent reduces the time needed to prepare a decision, while the administrator retains control over the change. Sales and commercial information management An agent can prepare a briefing before a meeting by bringing together information from the CRM, proposals, correspondence and notes, then highlighting open points and suggested next steps. After the meeting, it can create a summary, propose data updates and prepare tasks for the team. These are extensions of scenarios already familiar from everyday work with generative AI. Read more about what the current generation of models helps teams achieve in our article: GPT-5.6 from OpenAI: capabilities and business applications. Why does an AI agent need a workflow? An AI agent can interpret information and suggest next steps, but the process should define the sequence of actions, required validations and the people responsible for approval. In a large organisation, this is what determines the repeatability and scalability of the solution. A process automation platform can act as a control layer: it triggers a task, provides the agent with the necessary context, receives the result, records the history and routes the case to the next stage. The agent then becomes part of a controlled workflow rather than operating as a separate tool outside the core process. This approach is relevant to document workflows, request handling, HR processes, procurement and administration. See how WEBCON BPS can support the digitalisation and control of business processes, and how TTMS delivers process automation. Four forms of human oversight of an AI agent Human-in-the-loop is a model of control embedded in the process—from reviewing recommendations to handling exceptions and making decisions with greater impact. In a mature solution, people can play several different roles. Approving an action when the agent has prepared a specific instruction, communication or system change. Selecting an option when the agent has presented several possible solutions and their consequences. Handling an exception when a case falls outside the agent’s rules, available data or permissions. Overseeing process quality by analysing errors, rejected recommendations, completion times and changing business needs. The most effective implementations use all four forms. The team does not manually review every standard operation, yet retains full control over actions with greater significance and over the direction in which the process evolves. It is also worth observing whether human approval genuinely improves process safety or simply moves a bottleneck elsewhere. If an approver nearly always accepts the agent’s proposals without changes and the cases are easy to reverse, the organisation can consider automating the selected step. If recommendations often require correction or the approver needs to return to source data, this indicates that the process rules, quality of knowledge or scope of the agent’s permissions need attention. When can an AI agent act automatically? Automation delivers the most value when a task is frequent, has a repeatable structure, relies on available data and leads to a clearly defined outcome. It is also important to ensure that execution can be verified and corrected when data or rules change. Good candidates include ticket classification, routing requests to the appropriate queue, completing data from approved sources, creating standard tasks, updating statuses and sending communications based on approved templates. Combining GPT models with an enterprise knowledge layer, integrations and security rules provides a significant advantage. This allows the solution to work with information available to a specific role, rather than with an unstructured collection of documents and conversations. When should an AI agent primarily provide advice? An advisory role is especially valuable in cases that require contextual assessment, interpretation of company policy, negotiation, an individual approach to a customer or decisions with significant financial and legal consequences. In these situations, the agent can gather facts, summarise documents, identify missing information, compare options and prepare the rationale for a recommendation. The person gains time for business judgement, while the decision remains grounded in the knowledge, experience and accountability appropriate to the role. This model is particularly useful for managers, compliance specialists, legal teams, strategic procurement, finance teams and teams responsible for key accounts. FAQ What is the difference between an AI agent and a chatbot? A chatbot primarily responds to questions in a conversation. An AI agent can also use approved tools, retrieve information from enterprise systems, follow workflow rules and complete defined process steps. Its value comes from combining language understanding with access to business context, permissions and a controlled process. Should every AI agent have human approval before taking action? No. The appropriate level of oversight depends on the impact and reversibility of the action. Low-risk, repeatable activities such as categorising tickets or sending a standard confirmation can be automated under defined rules. Actions affecting customers, contracts, finances, compliance or production systems should usually include approval or escalation to an authorised person. Can one AI agent operate at different levels of autonomy? Yes. Autonomy should be designed for individual actions rather than assigned to an entire solution. The same agent may classify a request automatically, prepare a response for approval and escalate an unusual case to an expert. This makes it possible to automate safely without treating every task in the same way. What information does an AI agent need to work reliably in an enterprise? An agent needs access to reliable and current knowledge sources, a clearly defined objective, appropriate permissions and rules for handling exceptions. It should also receive only the context relevant to the task and role. Workflows, validations and an auditable history of actions help ensure that its output can be reviewed and used consistently. How can a company start implementing GPT-powered AI agents? Start with one clearly defined process step that has measurable volume, repeatable inputs and a known outcome. Set the boundaries of the agent’s permissions, test it with standard and exceptional cases, and measure the effect on process time, quality and escalations. Once the team has evidence that the solution works reliably, its scope and autonomy can be expanded gradually.

Read
ChatGPT 5.6 in Practice: Initial Compliments and Disappointments

ChatGPT 5.6 in Practice: Initial Compliments and Disappointments

OpenAI rolled out GPT-5.6 in stages. It first appeared in limited test access for selected partners. Access to ChatGPT 5.6 reached Europe, including Poland, gradually, so only recently have teams been able to test the model in everyday work. Expectations are high. In the second half of 2026, businesses expect language models to handle multi-step tasks and work with extensive context. Ease of use matters too. GPT’s interface has undergone a major redesign. Has it improved the user experience and the quality of responses? This article explores that question, as well as: which business processes ChatGPT 5.6 can support by improving productivity and the quality of working materials, how to plan an AI pilot in your organisation, measure results and maintain quality control, which limitations of ChatGPT 5.6 to consider before a wider rollout, how to establish a shared standard for prompts and output validation across the team, what early users think about working with ChatGPT 5.6. If you are looking for a full overview of the changes, pricing, models and capabilities of GPT-5.6, see our article GPT-5.6 from OpenAI: what has changed, pricing, capabilities and business applications. ChatGPT 5.6: our first impressions and early industry feedback Early expert reviews focus primarily on context handling. Reviewers note that when working with substantial material that goes through multiple rounds of edits, ChatGPT 5.6 is better at keeping the task on track. Most of us have experienced earlier OpenAI models losing their “bearing”. On top of that, the model itself encouraged endless revisions, which could pull the material away from the original intent of the prompt. GPT 5.5 had an irritating habit of suggesting more and more variations. Almost every response ended with a clickbait-style suggestion along the lines of: “If you want, I can help you add two elements that will create a wow effect and give the text around 50% more SEO power.” As a result, instead of closing the topic, we were drawn into the model’s endless doubts: could the material really not be improved further? GPT 5.6 is no less capable than the older model, but it finally respects what matters most: the intent behind the prompt and our time. Kajetan Terlecki SEO Specialist, TTMS Another recurring observation concerns the quality of the first draft—the material GPT produces after the first prompt. Reviewers emphasise that the model’s draft is usually well structured and much closer to a final version than it was with GPT 5.5. It is not a perfect ten yet, but a solid eight. In other words, a final version may be within reach after a relatively short time. With earlier GPT models, the “brainstorming” phase took much longer. The third—and most immediately noticeable—area is the way we use the tool, which we can simply call the “interface”. It is admittedly quite complex. Beyond writing a prompt, users must make a series of decisions: which workspace should I choose: Chat or Work? which model best fits my request: Luna, Terra or the most advanced Sol? Or is the older GPT 5.5 enough? does the task require Deep Research? how much effort should the model put into the task: low, medium, high, very high, max or ultra? should I use Turbo mode and generate a response 50% faster at the cost of higher token use? If we add the almost endless range of available plugins, writing the prompt turns out to be only half the work required to get a useful result. I would welcome an automatic mechanism that reads the prompt and selects the right settings on its own. One that uses a sufficiently capable GPT model without wasting tokens when they are not needed. How do you navigate all this? We have outlined a suggested configuration here, including which modes to use for different types of tasks. Where does GPT 5.6 outperform the previous version? 1. GPT 5.6 is better at preserving document layout and formatting The previous version of GPT had something of a goldfish memory. You could also compare it to a short blanket: pull it over one part, and another is left exposed. When we asked the model to update data in a document it had generated, it produced a factually correct response, but one that no longer followed the original format. It might use a different heading hierarchy, rearrange the information or omit elements that are essential for the company. GPT 5.6 is much better at preserving the structure of reference material. OpenAI illustrated the difference in materials introducing GPT-5.6. The company placed three slides side by side: the reference file, the GPT-5.5 output and the GPT-5.6 output. The task was to update figures in a presentation while retaining the original template. In the comparison, GPT-5.5 omitted some template elements, while GPT-5.6 preserved the slide structure more faithfully: layout, typography, spacing, colours and recurring template elements. OpenAI states that GPT-5.6 can also interpret rules saved in the slide template, including the Slide Master. In practice, this matters when a presentation needs to retain not only its colours and fonts, but also defined layouts, spacing and mandatory components. 2. GPT-5.6 moves beyond the chat window GPT-5.6 shows its greatest potential when it works not only with a single instruction, but also with files and tools made available by the user. It can then move quickly through a task: from gathering the materials to preparing a first draft. The new GPT model can identify related files in a project folder, flag places that need updating and prepare working versions of documents. There is a catch: the process still needs human oversight. Someone must check whether GPT found all the relevant files, understood the context correctly and left unchanged the elements that were meant to remain unchanged. Still, instead of manually digging through documents, the team starts with a list prepared by the model. 3. From an idea to a version you can show the team Experts testing GPT 5.6 point out that the first version of a simple application, dashboard or website is now more often suitable for showing to a team and collecting specific feedback. It is somewhat like an MVP: good enough to test an idea, present it to the team and gather initial comments. A product owner can see the whole process, a designer can assess the layout and usability, and a developer can spot technical constraints sooner. This does not mean that GPT-5.6 creates a finished product. The initial prototype still needs to be assessed for security, quality and architecture. The difference is concrete, however: the team can evaluate an actual solution earlier, rather than debating assumptions alone. 4. GPT 5.6: “I don’t know” — is this the end of answers given for the sake of answering? We all know the old classified ad: “Encyclopaedia Britannica, 40 volumes for sale. I got married a week ago, so I no longer need it. My wife knows everything better.” The know-it-all syndrome is a nuisance not only in old marriage jokes, but also for people who work with language models every day. GPT often lacks the information needed to give a reliable answer. GPT-5.5, like earlier versions, would rather provide an incorrect—yet convincing-sounding—answer than admit it did not know. What about the new version? The change is visible at first glance, even though it is hard to capture in a benchmark and easy to appreciate in day-to-day work. Our first days of working with the two most advanced models, Terra and Sol, suggest that GPT 5.6 is more likely to say “I don’t know”, “I don’t have enough data” or “I could not find anything else on this topic”. People still need to add or verify information manually, but this reduces the risk of an embarrassing error in material prepared for a client, the board or a project team. Before you give GPT-5.6 an important task: what to watch out for in early testing 1. A working prototype is not yet a finished product GPT-5.6 can prepare a website, dashboard or simple application that can be launched and shown to the team. This is a major step forward, particularly when testing an idea. The tests also reveal the other side: elements can become misaligned, interactions do not always work as intended, and visual details still require refinement. The first version can be an excellent starting point, but it should not automatically be sent to clients or other external audiences. Before treating it as finished, we need testing, a security assessment and, in some cases, a developer’s review. 2. The new Work environment can still be frustrating Model quality is one thing. The way we use it in practice is another. One reviewer pointed out that, in Work, it was difficult to access generated files and open a preview of the finished result. Others criticised the number of settings—discussed earlier in this article—as well as the unclear distinction between Chat, Work and Codex. GPT-5.6 may complete a task correctly, while the working environment still makes it difficult to retrieve or review the result. It is worth testing the entire process, not only the quality of the response in the chat window. 3. GPT needs clear boundaries One reviewer tested how GPT-5.6 would handle a complex mathematical problem. The model produced correct parts of the solution, but surrounded them with definitions, digressions and comments that added little value. Only after the instruction was made more specific did it produce a useful result. The same applies in a business context. We should not leave the model too much room for interpretation. It is better to state the expected result directly: “Prepare a one-page summary. Include the decision, three arguments, risks, missing information and next steps.” GPT then has fewer opportunities to pad the topic with peripheral content. 4. GPT can still be wrong The fact that GPT-5.6 appears more likely to signal that it lacks data or a basis for drawing a conclusion does not mean it is free from hallucinations. Luna, Terra and Sol—with Sol seemingly the least prone to this—can still provide an incorrect date, number, source or conclusion without batting an eyelid. The rule to “check after AI” still applies and will likely remain relevant for many future GPT releases. 5. Start with one problem, not a large system Once GPT-5.6 has access to files, a browser and company tools, it is easy to imagine a system that instantly organises the inbox, analyses team communication, updates the CRM and writes responses to clients. This vision can quickly turn into a project larger than the problem it was meant to solve. One expert working with an extensive Codex environment recommends starting with a single, repeatable task. It might be preparing a meeting summary, gathering open project issues or updating an offer after data changes. Only once the team sees measurable results and understands the tool’s limitations is it worth adding further automations. How should you run your first ChatGPT 5.6 test in the company? A pilot should answer one straightforward question: does GPT-5.6 genuinely improve a selected stage of work, and does the benefit justify the time, cost and additional quality control? The first test should not begin with building an extensive automation system. It is better to choose one repeatable task that currently takes up the team’s time and has a clearly defined outcome. This might be a meeting summary, a brief or a status report. What matters is that the team knows which materials it provides to the model, what result it expects and who reviews the final document. Before starting the pilot, answer five questions: Choose one process: for example, preparing meeting summaries, sales briefs or materials for project decisions. Set a baseline: measure the time needed to prepare the material, the number of revisions, the number of people involved and the most common errors. Prepare a shared prompt: use the same input materials and clearly describe the outcome the team expects. Assign expert review: nominate a person who will verify the facts, assess quality and approve the result before it is used further. Assess the outcome: compare time, the number of iterations, completeness of the material and the usefulness of the result for the next stage of the process. Pilot element Question for the team Process Which stage of work do we want to shorten or organise? Outcome What should be produced: a brief, decision list, analysis, recommendation or communication draft? Data Which materials are needed, and can they be used in the selected AI environment? Quality control Who confirms the facts, completeness and alignment of the material with the process? Metric How will we compare working time, the number of revisions and the usefulness of the result? After a few attempts, it becomes easier to assess whether the model is genuinely helping. Compare the time needed to prepare the material, the number of revisions and the effort required to verify the result. Only then decide whether to extend the pilot to further tasks. Three processes worth starting with 1. Summaries after client meetings The model can organise notes, gather decisions, identify open questions and prepare a list of next steps. The team confirms the arrangements and assigns task owners. This helps them move from discussion to action more quickly. 2. A brief for a sales conversation Based on selected sales materials, previous arrangements and public information about the company, GPT-5.6 can prepare a brief, discovery questions and a list of topics that require clarification. The salesperson remains responsible for the client relationship and decisions regarding the offer. 3. A status report for the project team The model can organise information about progress, blockers, risks and planned actions. The project owner confirms that the information is up to date before the report is shared further. This reduces the time the team spends manually consolidating data from several sources. How do you embed AI in a business process? After the pilot, it becomes clear whether ChatGPT 5.6 genuinely shortens the preparation of materials, reduces the number of revisions and helps the team move more quickly to the next stage of work. It also reveals where the model needs a better brief, access to data or expert oversight. Proven use cases can then be extended to other processes. At this stage, it is worth addressing data security, integration with existing tools, output quality and a clear division of responsibilities. These factors determine whether AI becomes lasting support for the organisation. At TTMS, we help organisations identify processes where automation and AI create business value. We then design solutions tailored to their data, regulatory requirements and ways of working. We combine engineering experience with a responsible approach to AI governance, confirmed by ISO/IEC 42001 certification. Let’s discuss the processes AI could support in your organisation. FAQ How do you choose a process for your first ChatGPT 5.6 test? The best candidate is a repeatable process that requires gathering several pieces of information and producing a predictable result. Examples include meeting summaries, sales briefs, status reports and document analysis. The team should know the current turnaround time and typical issues, as these provide the baseline for assessing the test. Start with one process and expand the use of AI only after evaluating the outcome. How do you measure the business value of ChatGPT 5.6? During a pilot, measure the time needed to prepare the first version of the material, the number of revisions before approval, the completeness of the output and the expert time required for verification. It is also useful to track metrics related to the next stage of the process – for example, faster meeting preparation, a shorter time to close agreed actions or fewer missing details in a report. This data helps assess team productivity based on actual results and supports decisions about integrating AI into further processes. What data should you prepare for working with ChatGPT 5.6? The model produces better results when the team provides current, well-organised source materials. Before starting, identify which documents take priority, which data must remain unchanged and how unverified information should be marked. The organisation should also define which data can be shared in the chosen AI environment. For personal, financial and confidential data, access rules, retention and compliance are essential. How do you maintain human oversight of the model’s work? Human oversight should be part of the process from the start. The process owner defines the task scope, an expert verifies facts and alignment with requirements, and an authorised person approves external actions. This division of responsibilities is particularly important for client communication, publications, data changes in systems and materials with legal or financial implications. It allows the team to use automation while retaining responsibility for the outcome. Where can I find information about GPT-5.6 pricing, models and capabilities? We have covered the changes in GPT-5.6, pricing, the Sol, Terra and Luna models, and business applications in a separate article: GPT-5.6 from OpenAI: what has changed, pricing, capabilities and business applications. This article focuses on the practical use of ChatGPT 5.6 in team workflows, early user experiences and how to run an AI pilot in an organisation.

Read
Best AI Governance Solutions for Regulated Industries in 2026

Best AI Governance Solutions for Regulated Industries in 2026

In 2026, regulated enterprises cannot scale AI without governance. Every AI system that affects business decisions, customer data or operational risk needs clear ownership, documented controls, human oversight and post-deployment monitoring. The pressure is no longer theoretical. The EU AI Act is already in force, GPAI obligations have started to apply, transparency requirements are becoming operational, and sector-specific expectations around digital resilience, model risk and data protection remain active in finance, healthcare, energy, life sciences, public sector and other regulated environments. At the same time, ISO/IEC 42001 has become one of the clearest management-system standards for turning AI governance from policy language into operating reality. TTMS Expert Insight “In regulated industries, AI governance cannot remain a policy document. It has to become part of how AI systems are designed, delivered, monitored and improved every day.” Adam Kaczmarczyk Chief Operating Officer, TTMS That is why the search for the best AI governance solutions for enterprises 2026 should not end with a shallow top-10 ranking. Regulated organizations do not need software alone. They need an operating model, clear controls, audit-ready evidence and implementation discipline. The best AI governance solutions help enterprises connect policy, technology, risk management and daily business operations. In practice, this means comparing different categories of enterprise AI governance solutions: broad governance suites such as IBM watsonx.governance, Credo AI and Dataiku Govern; ecosystem-based platforms such as Microsoft Purview and Google’s Gemini Enterprise Agent Platform; and specialist observability or runtime-control vendors such as Fiddler AI and Arthur AI. Open-source projects also matter, especially for technical teams, but in regulated environments they usually work best as components of a wider governance architecture rather than complete governance systems. 1. What Are AI Governance Solutions? AI governance solutions are technologies, frameworks and operating models that help organizations manage AI responsibly throughout its lifecycle. They support activities such as AI inventory, risk assessment, documentation, monitoring, human oversight and regulatory compliance. Unlike traditional IT governance, AI governance focuses on how models, applications and AI agents are developed, deployed, monitored and retired while maintaining transparency, accountability and regulatory compliance. 2. Why AI Governance Is Becoming a Board-Level Priority The EU AI Act is the most important regulatory starting point for many European organizations. It introduces a risk-based approach to AI and places particular attention on use cases such as critical infrastructure, education, employment, essential services including credit scoring, biometrics, law enforcement, migration and the administration of justice. For high-risk AI systems, the required governance elements closely match what modern AI governance solutions are designed to support: risk assessment and mitigation, dataset quality, logging for traceability, technical documentation, clear information for deployers, human oversight, robustness, cybersecurity and accuracy. Organizations should also be aware that AI Act implementation is not a single deadline. Different obligations enter into force at different stages, depending on the type of AI system, sector and use case. This makes governance readiness essential. Enterprises need to prepare documentation, supplier oversight, monitoring processes and operating-model maturity before compliance pressure becomes urgent. This is why regulated industries are the natural audience for AI applications governance solutions and enterprise AI governance solutions. Financial services face overlapping expectations from the AI Act, model-risk management and digital operational resilience. In Europe, DORA has applied since January 2025 and covers ICT risk management, third-party risk, resilience testing, incident reporting and oversight of critical providers. Regulatory Readiness AI Act compliance is not a single deadline. It is a staged journey that requires governance readiness across data, models, vendors and business processes. Risk-Based Approach Classify AI systems based on their use case, business impact and regulatory exposure. High-Risk Controls Prepare documentation, logging, human oversight and cybersecurity controls. Sector-Specific Requirements Align AI governance with DORA, model risk management and data protection requirements. Third-Party AI Govern external LLMs and SaaS AI tools through vendor oversight and output validation. The same logic extends beyond banking. Healthcare, life sciences, insurance, utilities, energy, public sector and HR-intensive organizations all need mature solutions for AI governance, even when they are not training frontier models themselves. Companies using external LLMs or SaaS-based AI still need oversight, documentation, vendor accountability, data controls and human review. 3. Who Needs AI Governance? Any organization using AI in business-critical, regulated, customer-facing or high-impact processes needs AI governance. This includes companies building their own AI systems and companies using third-party tools embedded in daily operations. AI governance is especially important when AI influences decisions about people, money, health, safety, legal rights, employment, access to services or regulated business processes. In these contexts, governance is not only about avoiding mistakes. It is about proving that decisions, data flows, models, vendors and controls are managed responsibly. 4. Which Industries Require AI Governance Most? AI governance is most urgent in regulated industries where AI decisions can create legal, financial, operational or reputational risk. These include: financial services and insurance, healthcare and life sciences, energy and utilities, public sector and administration, transport and critical infrastructure, legal services, HR and recruitment, manufacturing and safety-critical industries. In these sectors, AI governance is becoming part of broader enterprise risk management. The key question is no longer whether AI should be governed, but how to make AI controls auditable across data, models, applications, vendors and operations. 5. What Regulations Affect AI Governance? Several regulatory and standards-based frameworks influence how organizations govern AI in 2026. The EU AI Act is the central framework for AI systems in the European Union. DORA affects digital operational resilience in the financial sector. Model-risk management expectations remain important for financial institutions. Data protection laws continue to shape how personal data can be used in AI systems. ISO/IEC 42001 is also becoming highly relevant because it gives organizations a structured way to manage AI through a formal AI management system. It applies not only to organizations developing AI-based products and services, but also to those using AI in their operations. For regulated enterprises, the practical task is to translate these requirements into everyday controls: ownership, documentation, risk classification, data quality, human oversight, monitoring, vendor assessment and audit evidence. AI Governance Framework Snapshot EU AI Act Risk-based legal framework for AI systems in the European Union. ISO/IEC 42001 Management system standard for governing AI across the organization. DORA Digital operational resilience requirements for financial institutions. Data protection laws Rules governing personal data processing in AI systems. 6. How Do AI Governance Platforms Work? Most top AI governance solutions companies now converge around a similar lifecycle. A governance platform typically starts with inventory: what AI systems exist, who owns them, what data they touch, what business purpose they serve and which regulations apply. From there, the platform maps policies to controls, supports validation and approvals, collects evidence and continues after deployment with monitoring, alerts, incident handling, retraining or re-approval workflows and audit reporting. Buyers searching for AI-powered data governance solutions, automated AI governance solutions and data governance solutions for AI systems are usually looking for the same thing: a repeatable evidence trail from use-case intake to runtime control. Key Takeaway The best AI governance platforms do not simply monitor models. They create an auditable chain of evidence across the entire AI lifecycle. 01 Data Source, quality and permissions 02 Models Evaluation, testing and versioning 03 AI Agents Roles, actions and permissions 04 Business Owners Accountability and approvals 05 Regulatory Controls Policies, evidence and audit trails 06 Operational Monitoring Alerts, incidents and continuous review 6. Seven Capabilities Every Enterprise AI Governance Solution Should Provide 1. Enterprise-Wide AI Inventory and Ownership The platform should discover and catalog models, applications and agents, including shadow AI. Enterprises need to know what exists, who owns it, what data it uses and what business risk it creates. 2. Risk Classification and Control Mapping A serious governance platform should classify AI systems by risk and map those risks to internal policies, regulatory obligations and control requirements. This is essential for regulated industries and aligns with the risk-based logic of the EU AI Act. 3. Data Governance, Provenance and Traceability High-quality data, logging, documentation and traceability are not optional in regulated AI. Strong AI-powered data governance solutions help organizations understand where data comes from, how it is used and whether it is appropriate for a specific AI use case. 4. Evaluation, Testing and Runtime Monitoring AI systems should be tested before deployment and monitored after deployment. This includes checks for drift, bias, performance degradation, unsafe outputs, security issues and unexpected behaviour. 5. Human Oversight, Approvals and Escalation Regulated organizations need clear approval workflows, sign-offs, separation of duties and escalation paths. The best governance systems do not remove human responsibility. They make it visible and auditable. 6. Explainability, Audit Evidence and Reporting Strong governance solutions for AI model transparency turn governance activity into documentation, reports, evidence trails and decision history. This is where broader AI transparency and governance solutions become operational rather than theoretical. 7. Third-Party and Agent Governance AI governance can no longer stop at internal models. Enterprises increasingly rely on third-party models, SaaS AI tools and AI agents. This creates new requirements around vendor oversight, permissions, runtime behaviour, logging and intervention paths. AI Governance Lifecycle for Regulated Enterprises Most mature AI governance programs follow a repeatable lifecycle that connects business ownership, regulatory mapping, technical validation and audit evidence. Use case intake – identify the business purpose, expected value, affected users and potential risk. AI inventory and ownership – register the AI system, assign an accountable owner and document the systems, data and vendors involved. Risk classification – assess regulatory exposure, business impact, data sensitivity and potential harm. Data and provenance review – verify data quality, source, permissions, security and suitability for the AI use case. Model or agent evaluation – test performance, robustness, bias, explainability, safety and alignment with business requirements. Human approval – define approval workflows, escalation paths and human oversight before deployment. Deployment control – release the AI system with documented controls, access rules and monitoring requirements. Runtime monitoring – track performance, drift, errors, incidents, user feedback and unexpected behaviour. Corrective action – manage incidents, exceptions, retraining, configuration changes or suspension when needed. Periodic review – reassess the system regularly and decide whether to continue, update, retrain or retire it. Audit evidence – maintain documentation, logs, approvals and control records for compliance and internal assurance. 10. Comparative Landscape of Leading AI Governance Platforms The field of top AI governance solutions companies is broad enough that a single-winner ranking is misleading. Different products solve different parts of the governance challenge. The table below is not a ranking. It is a role-based comparison for regulated buyers. Solution Best for Main strengths Limitations Microsoft Purview Microsoft-centric enterprises needing strong data security, compliance, audit and catalog foundations Strong fit for AI-powered data governance solutions, including data governance, audit, information protection, compliance and lifecycle management Less of a dedicated standalone AI risk suite; works best as a control foundation inside a broader Microsoft architecture IBM watsonx.governance Large regulated enterprises needing policy-to-control mapping across hybrid environments Strong governance graph, policy mapping, continuous reporting, regulatory content and AI/GRC integration Can be heavyweight for organizations looking for a narrow or lightweight use case Google Gemini Enterprise Agent Platform Google Cloud users building models and agents inside one engineering stack Strong model evaluation, registry, monitoring, secure development and governed enterprise-agent capabilities More platform-centric than governance-program-centric; may require additional compliance orchestration Credo AI Enterprises wanting centralized AI inventory, risk intelligence and regulatory mapping Strong registry, shadow-AI discovery, policy packs, evidence generation and governance across models, agents and applications Some teams may still pair it with separate model platforms or observability tools Dataiku Govern Organizations wanting governance embedded into the AI delivery workflow Strong workflows, registries, sign-off rules, audit timelines, LLM registry and growing agent-management capabilities Best fit when Dataiku is already part of the AI operating model Fiddler AI Runtime-heavy environments focused on monitoring, guardrails and observability Strong for continuous evaluation, root-cause visibility, inline enforcement and agentic monitoring More specialized around observability and runtime control than full enterprise management-system governance Arthur AI Teams prioritizing agent discovery, evaluation, observability and guardrails Good coverage of agent discovery, performance evaluation, built-in guardrails and model-agnostic support Less public emphasis on regulatory content libraries and formal enterprise compliance workflows MLflow Engineering-led teams needing open-source observability, evaluations, registries and model management Useful open-source backbone for custom AI governance stacks Not an out-of-the-box regulatory governance suite Evidently Teams needing open-source testing, monitoring and dashboards Strong for evaluating, testing and monitoring ML and LLM systems Not a complete governance operating system for policy, accountability or regulatory workflows Giskard LLM and agent teams focused on testing, red-teaming and evaluation Useful for LLM and agent safety, security and validation workflows Not a full enterprise governance suite with broad policy packs and formal approval routing AIF360 / Fairlearn Organizations needing open-source fairness assessment and bias mitigation Mature tooling for detecting and mitigating bias Best treated as components inside a wider governance design, not as end-to-end solutions for AI governance The practical pattern is clear. Platforms such as IBM, Credo AI and Dataiku are closer to end-to-end governance layers. Microsoft Purview and Google’s platform are powerful when governance is tightly linked to data estates and cloud engineering. Fiddler and Arthur are strongest where runtime performance, decision lineage, agent control and guardrails matter most. Open-source projects are indispensable for cost-effective experimentation and specialized controls, but they usually need architectural composition before they resemble full enterprise AI governance solutions. 11. Open-Source vs Commercial AI Governance Tools Organizations considering the best open-source AI governance solutions 2026 should take a toolkit view rather than look for one universal platform. Open-source is strong in technical subdomains: fairness and bias mitigation with AIF360 and Fairlearn, observability and drift monitoring with Evidently, evaluation and testing for LLM agents with Giskard, and AI engineering workflows with MLflow. These tools can be highly valuable, especially for engineering-led organizations. However, they are usually not full business governance systems. They do not, by themselves, deliver the full mix of regulatory mapping, approval workflows, ownership assignment, cross-functional reporting and audit-ready evidence that commercial governance suites emphasize. Commercial tools, by contrast, usually win on speed to governance. They package inventory, workflows, policy libraries, integrations, alerts, evidence capture and executive reporting in ways that better serve compliance, risk, procurement and audit teams. For regulated enterprises, the right answer is often hybrid: commercial governance platforms for enterprise control and reporting, supported by open-source tools for specific technical evaluations, monitoring or fairness checks. 13. Why Agentic AI Needs Separate Governance AI agents introduce a new governance challenge. Unlike traditional AI models that generate an output for a human to review, agents can plan, call tools, access systems, trigger workflows and perform multi-step actions. This changes the risk profile. Enterprises need enterprise AI agent governance solutions that can define what an agent is allowed to do, which systems it can access, what data it can use, when a human must approve an action and how every step is logged. Governance must cover the agent’s role, permissions, model behaviour, tool access, output quality, runtime monitoring and intervention paths. This is why agent governance should not be treated as a footnote to model governance. It requires its own inventory, approval workflows, control design, monitoring and incident response model. AI Agent Governance Checklist Every enterprise deploying AI agents should be able to answer these questions before production. ✓ What systems can it access? ✓ What data is the agent allowed to access? ✓ What actions is the agent allowed to perform? ✓ When is human approval required? ✓ Is every action logged? ✓ Can the agent be stopped immediately? ✓ Who is accountable for the agent? Organizations that cannot answer these questions before deployment will struggle to demonstrate effective governance once AI agents begin interacting with enterprise systems and business processes. 14. How to Choose the Right AI Governance Solution The best buying logic for regulated enterprises starts with the problem, not the vendor demo. If the main challenge is data sprawl, sensitive information control, audit and compliance across Microsoft environments, Microsoft Purview may be a strong foundation. If the priority is enterprise-wide policy management and regulatory mapping, IBM watsonx.governance, Credo AI or Dataiku Govern may be more relevant. If the business needs runtime quality control, observability, guardrails and agent monitoring, Fiddler AI or Arthur AI may become stronger candidates. If the organization is engineering-heavy and prepared to design its own operating model, open-source stacks based on MLflow, Evidently, Giskard and fairness libraries can be powerful. Second, test the platform against the regulatory footprint, not only the presentation. Regulated buyers should ask whether the solution supports risk classification, data quality and provenance, audit evidence, human oversight, third-party governance and post-deployment monitoring. Third, check whether the platform can support governance across the full AI estate: models, applications, agents, vendors, data pipelines and business processes. AI governance that only works for one model or one team will not scale across a regulated enterprise. 15. Why AI Governance Is More Than Software AI governance software can support discovery, workflows, evidence and monitoring, but it cannot define accountability on its own. Regulated organizations need a governance operating model that connects business owners, compliance, legal, data teams, security, IT, procurement and executive leadership. This is where AI governance consulting & solutions become essential. The platform is only one part of the answer. Organizations also need to define what AI use cases are allowed, how risks are classified, who approves deployment, what evidence is required, how vendors are assessed, how incidents are handled and how governance evolves as AI systems change. Without this operating model, even a strong platform becomes another dashboard. With the right governance framework, AI can move from pilots to production in a way that is controlled, auditable and aligned with business goals. 16. TTMS Project Insight: Governance Starts Before the Model One lesson we have seen repeatedly in client projects is that governance challenges rarely begin with the AI model itself. They usually start much earlier: with the quality of source documents, inconsistent business processes, fragmented knowledge and unclear ownership of information. In one TTMS project for a law firm, we developed an AI solution supporting court document analysis. While selecting the right language model was important, the biggest implementation effort focused on preparing trusted legal content, defining document workflows, validating AI-generated outputs and ensuring that lawyers remained in control of final decisions. Governance became an integral part of the solution rather than an additional compliance layer. The same pattern appears across regulated industries. Organizations often discover that successful AI adoption depends less on choosing the “best” model and more on establishing reliable governance around data, processes and human oversight from the very beginning. In our experience, organizations rarely struggle because they chose the wrong AI model. More often, they struggle because they underestimated the governance needed around it. Read more about this project in our AI implementation for court document analysis case study. You can also explore more examples in the TTMS case studies library. 17. How TTMS Helps Regulated Enterprises Govern AI TTMS supports organizations that need to move from AI ambition to governed AI implementation. As an AI consulting and strategy partner, TTMS helps regulated enterprises assess AI risk, design governance frameworks, select suitable governance architecture and operationalize controls across data, models, applications, vendors and agents. The company’s approach is strengthened by its ISO/IEC 42001-certified AI Management System. TTMS states that this system governs both internal and external AI-related projects delivered under the TTMS brand. This matters because AI governance is not only a client advisory topic. It is also a way of working that must be reflected in project delivery, documentation, risk management and operational oversight. For organizations using third-party AI tools, this is especially important. Governance is still required even when the AI model is not built in-house. Enterprises need to understand how external tools use data, how outputs are reviewed, what risks are introduced, which controls are required and how accountability is maintained. TTMS helps clients approach AI governance as a practical implementation challenge rather than a documentation exercise. The goal is not to slow innovation down, but to make AI adoption safer, more scalable and easier to defend in regulated environments. 18. From AI Governance Strategy to Practical Business Solutions Choosing the right AI governance platform is only one part of building a successful AI strategy. Organizations also need practical governance frameworks, clear policies, evidence workflows, vendor assessment, risk classification and implementation expertise that connects technology with business and regulatory requirements. At TTMS, we combine AI governance consulting & solutions with the development of secure, enterprise-ready AI products. Rather than offering a single generic AI platform, TTMS develops specialized solutions for individual business processes, allowing organizations to combine practical AI adoption with governance, security and regulatory compliance. This approach helps enterprises move from strategy to implementation: from selecting enterprise AI governance solutions and defining controls to deploying AI tools that support real operational needs in legal, document analysis, e-learning, knowledge management, localisation, AML, recruitment and software testing. AI4Legal helps legal teams analyse court documents, generate contracts and process hearing transcripts while maintaining full control over sensitive legal information. AI4Content enables secure document analysis and knowledge extraction, generating structured summaries and reports in controlled cloud or on-premise environments. AI4E-learning transforms internal documentation into complete e-learning courses, helping organizations scale AI literacy and workforce development. AI4Knowledge provides employees with governed access to organizational knowledge, procedures and internal documentation through conversational AI. AI4Localisation automates multilingual content translation while preserving terminology consistency and industry-specific language. AML Track supports anti-money laundering processes through automated screening, reporting and fully auditable compliance workflows. AI4Hire assists HR teams with CV analysis, candidate matching and resource allocation using transparent,>QATANA improves software quality by automating test management and AI-assisted test case generation in secure enterprise environments. All of these solutions are developed and delivered within TTMS’s AI Management System aligned with ISO/IEC 42001. This means clients benefit not only from innovative AI technology but also from established governance practices covering risk management, documentation, human oversight, security and regulatory compliance throughout the entire AI lifecycle. Whether your organization is evaluating enterprise AI governance solutions, looking for AI governance consulting & solutions, or planning to deploy AI in a regulated environment, TTMS helps turn governance into a practical business capability that enables innovation instead of slowing it down. FAQ What are the best AI governance solutions? There is no single universal winner. The best AI governance solutions depend on the enterprise problem. IBM watsonx.governance, Credo AI and Dataiku Govern are among the strongest broad governance suites. Microsoft Purview is highly relevant when data governance, compliance and Microsoft-stack integration dominate. Google’s Gemini Enterprise Agent Platform is strong for teams building governed agents and models in Google Cloud. Fiddler AI and Arthur AI can be excellent where runtime observability, agent control and guardrails are the priority. Open-source stacks can also be valuable, but usually as components rather than complete enterprise governance systems. What are the best open-source AI governance solutions in 2026? For buyers asking about the best open-source AI governance solutions 2026, the strongest answer is a toolkit view. MLflow is a broad open-source AI engineering base. Evidently is strong in testing and monitoring. Giskard is especially relevant for LLM and agent evaluation. AIF360 and Fairlearn are useful for fairness analysis and bias mitigation. However, most regulated enterprises will still need additional workflow, policy, reporting and audit layers on top. Can AI governance be automated? Yes, but only partially. Inventory, control mapping, evidence collection, recurring checks, continuous evaluations, alerts and parts of reporting can be automated effectively. Accountability decisions, material risk acceptance, exceptions and final approvals should remain under human oversight. The best automated AI governance solutions support governance teams instead of replacing them. Do organizations need ISO/IEC 42001 if they only use third-party AI tools? Certification is not always mandatory, but the standard is highly relevant for organizations using AI in regulated, customer-facing, high-impact or procurement-sensitive contexts. ISO/IEC 42001 is designed for organizations providing or using AI-based products and services. Even companies relying on external AI tools still need oversight, documentation, vendor accountability, data controls, risk assessment and human review. How should enterprises govern agentic AI? Enterprises should treat AI agents as a higher-governance category than ordinary chatbots. Agents need inventory, role and permission boundaries, model evaluation, action controls, logging, runtime monitoring and intervention paths for unsafe or off-policy behaviour. This is why the market is shifting toward enterprise AI agent governance solutions and why agent governance should be designed separately from traditional model governance. What Do Analyst Ratings Say About AI Governance Solutions? Publicly available best AI governance solutions analyst ratings should be treated carefully because many detailed comparisons from Gartner, Forrester and IDC sit behind paywalls. Still, public vendor disclosures and analyst mentions show a clear direction of travel. The market is rewarding platforms that provide centralized AI inventory, risk management, continuous monitoring, policy enforcement, evidence generation and agent/runtime governance. This is also why the search intent behind best AI governance solutions risk management 2026 is shifting away from one-time ethics checklists and toward continuous control planes. For regulated enterprises, this is the right direction. AI governance is converging with operational resilience, cybersecurity, data governance and enterprise risk management.

Read
1
240