GPT-6 Cyber: What OpenAI’s Cybersecurity Model Means for Business
OpenAI has introduced GPT-6 Cyber, an AI model designed for cybersecurity work. For companies managing customer portals, payment services, and internal platforms, potential benefits include faster vulnerability investigations, more frequent security reviews, and better support for engineers preparing and testing fixes. Its practical value will depend on whether it helps teams confirm vulnerabilities and test fixes while reducing the time specialists spend on those tasks. What Is GPT-6 Cyber? GPT-6 Cyber is positioned as OpenAI’s specialized model for cybersecurity work. Before its introduction, BusinessToday reported that selected customers were testing an alpha version through Daybreak Red. According to Investing.com, citing Fortune, OpenAI is also developing a separate product to help customers deploy GPT-6 Cyber securely, automate security tasks, and patch vulnerabilities. The product would give OpenAI greater oversight of how its models are used. For businesses, this raises practical questions about how the service would fit into existing security operations: which tasks it could carry out, which actions would require approval, and what information OpenAI would receive about its use. These details would influence both implementation costs and the suitability of the service for handling sensitive company systems. There is already a documented foundation for this approach. OpenAI describes GPT-5.6 Cyber as a model for approved users conducting advanced vulnerability research and security testing. It also offers Codex Security, an application security agent that helps teams investigate vulnerabilities and prepare fixes. These products show how OpenAI already supports security work; GPT-6 Cyber needs its own assessment of capabilities and performance. For a buyer, three parts of an AI security solution deserve separate attention: Part of the solution What it determines What the business should verify Model The analysis and reasoning available for a task Accuracy on relevant security cases, limitations, and cost Access program Who can use particular capabilities and under which conditions Eligibility, approval requirements, and permitted activities Application and integrations The data the system can inspect and the actions it can execute Supported tools, permissions, review controls, and audit records Before purchasing access, check how the solution will connect to the company’s code repositories, monitoring tools, and process for approving changes. Why Cybersecurity AI Matters to Business Leaders Security work competes for engineering capacity. Investigating a suspicious code change, checking a supplier’s advisory, and validating a patch all draw on people who also maintain applications and deliver new features. AI creates an opportunity to move some of the research and preparation into a repeatable workflow, allowing specialists to spend more time on judgment, validation, and decisions about production systems. Faster security work can also help companies keep their services running. A retailer may need to secure a checkout integration without disrupting sales. A manufacturer may need to understand the consequences of updating software connected to operational processes. A financial services company may need clear evidence of what was investigated and how a problem was resolved. In each case, useful automation must fit the business process surrounding the software. OpenAI’s broader investment shows the importance it assigns to this market. On September 3, 2026, the company announced a $1 billion commitment covering subsidized Daybreak access, training, technical support, and partnerships for organizations protecting essential services. Four Business Uses to Evaluate for GPT-6 Cyber The following scenarios are our analysis of where a cybersecurity model could create value. They are proposed evaluation areas whose suitability for GPT-6 Cyber requires testing. Each depends on the capabilities, tools, and permissions available in the eventual deployment. 1. Investigating Vulnerabilities in Business Applications A useful pilot could focus on a small set of applications with a clear business owner. The model would be given approved code and architecture information, then assessed on whether it helps explain suspected weaknesses and identify the evidence needed to confirm them. For example, a team might investigate whether a customer portal consistently enforces access permissions across related services. The business benefit would come from reducing the time needed to reach a defensible conclusion. A helpful result should identify affected components, explain the conditions under which the issue matters, and make it easier for an engineer to reproduce the problem in an authorized test environment. 2. Preparing and Testing Security Fixes Once a vulnerability is confirmed, the next task is to prepare a change that addresses it. A security AI workflow could help draft a patch, identify related code paths, and propose tests for both the security issue and normal application behavior. Existing Codex Security documentation already describes workflows for reviewing changes, validating findings, and preparing fixes, making this a concrete area for evaluating future model improvements. For the business, success means less engineering effort per validated fix while maintaining release quality. Engineers should still review the change and run appropriate tests before it reaches production. 3. Prioritizing Work Using Business Context Using an up-to-date list of systems, their owners, and supporting documentation, an AI system could help connect technical findings with business consequences. This could help the team decide which vulnerabilities to fix first and who should handle them. 4. Helping Teams Investigate Security Incidents Another use to test is building an incident timeline from security alerts and system logs, with internal procedures guiding the investigation. An evaluation could test whether the model links its statements to evidence, distinguishes observations from hypotheses, and identifies missing information. This scenario would require suitable integrations; it should not be assumed to be a built-in GPT-6 Cyber feature. A useful output could help the next analyst continue an investigation or help a service owner understand the affected business process. Decisions such as disabling accounts, blocking traffic, or isolating systems need explicitly assigned authority because they can interrupt legitimate activity. What Could This Look Like in a Customer Portal? Consider a company preparing a new version of a customer portal connected to its order management system. A security review flags a possible weakness in how the portal checks access to order details. In a controlled pilot, an AI system would receive the relevant code, a description of expected permissions, and test accounts containing fictional customer data. The team would assess whether it can help trace the affected logic, produce a clear explanation, and propose a test that demonstrates the problem. If the finding is confirmed, engineers could evaluate its suggested correction and additional regression tests, which check that existing functions still work after a change. The final result should be a reviewed change with evidence: what was wrong, what was modified, which tests passed, and who approved the release. This would require access to the relevant code, a test environment, and a process for assigning confirmed issues to engineers. How to Measure the Business Value of GPT-6 Cyber A pilot should compare the AI-assisted workflow with the team’s current process using comparable cases. Include confirmed vulnerabilities, issues previously dismissed as false alarms, and cases where the evidence is incomplete. Record analyst review time as well as model execution time. An answer produced quickly can still require substantial investigation. The following measures provide a practical basis for a decision: Measure What to record Why it matters Time to assess a suspected vulnerability Elapsed time and analyst effort needed to confirm or dismiss an issue Shows whether the workflow accelerates investigation Finding quality Confirmed findings, false alarms, and missed issues in a reference test set Reveals whether apparent productivity comes with additional errors Time to a tested and approved fix Time from confirmation to a tested, approved correction Connects AI assistance to remediation Engineering effort Hours spent reviewing, correcting, testing, and documenting outputs Makes supervision costs visible Total cost per resolved case Model use, tools, test environments, integration, and staff time Supports a realistic decision about scaling For example, if a pilot saves analyst time but generates extra work for developers, both effects belong in the assessment. If it uncovers more valid vulnerabilities, the organization also needs capacity to fix them. Access, Pricing, and Deployment Questions for Buyers OpenAI’s existing Daybreak documentation requires appropriate organization approval and project access. It distinguishes access to a program from access to a particular model. For GPT-6 Cyber, organizations should verify the applicable requirements directly, including whether access is limited to an evaluation or supports the planned production use. Commercial evaluation should cover pricing, usage limits, supported integrations, and the handling of company data. Teams should establish what source code, logs, and configuration information would be processed; where processing occurs; how long data is retained; and which contractual controls apply. Availability through a particular cloud provider or within an existing subscription should be confirmed for the specific offering. A pilot budget should include implementation and review costs alongside any model usage fees. How to Keep AI Security Work Under Control A business should define the scope of an AI security workflow as carefully as it defines the scope of an external security assessment. Specify the systems it may inspect, the information it may access, and the actions it may take. Separate permission to analyze a problem from permission to change a live service. OpenAI’s published Daybreak guidance recommends isolated environments, monitoring of agent actions, and enforced limits on authorized activity. For an initial business pilot, that supports a controlled test environment, narrowly scoped credentials, review of proposed changes, and records that allow the team to understand what happened. The operating process also needs an owner. Someone must review unresolved findings, decide when additional evidence is required, and stop a workflow that behaves unexpectedly. Training should cover how to challenge an AI-generated conclusion and how to recognize when the system lacks enough information to proceed. Where Businesses Should Start Start with one application and one recurring security task, such as investigating suspected vulnerabilities or testing fixes. Assign someone to review the results and agree how you will measure accuracy, time saved, and the effort required from engineers. A focused pilot can help establish whether the approach is useful enough to expand. Considering AI for your company’s security processes? Talk to TTMS about the task you want to improve, the systems involved, and your data requirements. Together, we can explore a suitable approach and define what a useful pilot should demonstrate. Discuss your AI use case with TTMS How is GPT-6 Cyber different from using a general-purpose AI model for cybersecurity? GPT-6 Cyber is described in published reports as a model focused on cybersecurity, with early testing through OpenAI’s Daybreak Red program. General-purpose models can also assist with security tasks, such as explaining code, reviewing documentation, and analyzing supplied findings. The question for a business is whether the specialist model produces more accurate, useful results on the tasks its team actually performs. Its name alone does not establish that advantage. A meaningful comparison should give both models equivalent information, tools, and time, then have security specialists assess the results. Published evaluations can help inform that comparison, but results from another model should not be attributed to GPT-6 Cyber. Can GPT-6 Cyber replace a cybersecurity team or an external security provider? A company should retain qualified people responsible for assessing risks, validating findings, and approving changes. A model’s analysis depends on the information and tools available to it, which may leave gaps in its understanding of the company’s systems. Security specialists also consider operational priorities, investigate ambiguous evidence, and coordinate the response when something goes wrong. AI assistance may reduce the effort needed for particular tasks, but that needs to be demonstrated in practice. For a business using an external security provider, a useful discussion is how the provider validates AI-generated work and whether any time savings improve the service. Responsibility for protecting systems and resolving problems should remain clearly assigned. Is GPT-6 Cyber useful for a company that mainly uses software from external vendors? Its usefulness would depend on what the company controls and what information it can access. A business using standard cloud applications may have limited visibility into the vendor’s underlying code. Its own responsibilities may include account permissions, configuration, integrations, and custom extensions. Those areas could provide relevant tasks for AI-assisted analysis, subject to the model’s verified capabilities and the available integrations. Before testing, establish which systems the company is authorized to assess and what requires the vendor’s involvement. Findings affecting the vendor’s product should be reported through its security process so that they can be investigated and addressed. Does an AI security review that finds no vulnerabilities mean an application is secure? No. A review can miss vulnerabilities because of incomplete information, limited test coverage, or errors in the analysis. A finding-free report should therefore describe what was examined, which tests were performed, and what remained outside the scope. For example, reviewing selected source files does not establish that the application’s production configuration and access permissions were also checked. The result should be considered alongside other security evidence, including testing and specialist review appropriate to the application’s risk. Record unresolved questions and limitations so that a clean report does not create unjustified confidence. How should a company assess a supplier offering “GPT-6 Cyber-powered” services? Ask the supplier to explain which model it uses, how it accesses that model, and which parts of the service rely on it. Request a demonstration using a representative, authorized task and ask to see the evidence supporting the findings. Establish who reviews the output, who implements corrections, and what happens when the analysis is wrong or incomplete. The service description should also explain how company data is handled and which actions require approval. If the supplier changes the underlying model, ask how it checks that the service still meets the agreed requirements. Evaluate the offer against clear deliverables, such as validated findings and tested fixes, with named people responsible for the work.
ReadRAG for Chatbots using CrewAI: Notes from a TTMS Tech Talk
Tech Talk is an internal series of technology sessions for TTMS employees, where we share knowledge and project experience. We demonstrate tried-and-tested tools, discuss challenges we have encountered and explain the solutions that have helped us in our work. Topics include artificial intelligence, data analytics, Salesforce, AEM and project management. Presentations are followed by time for questions, discussion and sharing ideas. During the session “RAG for Chatbots Using CrewAI”, held on 17 September, Jakub Kraśniewski, Senior AI Developer at TTMS, discussed improvements to a chatbot using a client’s documentation. He presented the challenges involved in preparing and retrieving information, the solution implemented and the approach to evaluating answer quality. The project involved a company in the education sector whose customers were preparing for a certification exam. The chatbot was intended to help them find information about registration, exam procedures, grading and appeals, thereby reducing the support team’s workload. It used several hundred pages of publicly available PDF documents, mainly in English. The team needed a way to retrieve relevant information from these materials while meeting a response time requirement of around 4 to 5 seconds. 1. How does RAG help a chatbot use company knowledge? Jakub began the presentation by explaining how RAG (Retrieval-Augmented Generation), a method of generating answers using retrieved source material, works. The system finds information relevant to the user’s question and passes it to a language model as context for the answer. In this project, the retrieved material consisted of passages from documentation describing exam rules and procedures. After extracting text from the documents, the system divides it into smaller chunks. An embedding model (an AI model that represents semantic features of text as numbers) converts these chunks into vectors stored in a database. The user’s question is processed in the same way. Comparing these representations allows the system to retrieve passages that are semantically related to the question. Jakub emphasised that the team is responsible for the quality of the material passed to the model. This involves checking whether the text was extracted correctly, whether the chunking preserved the necessary context and whether retrieval provides information useful for answering the question. The challenges the team encountered in the CrewAI-based solution demonstrated the importance of these steps. 2. What made information retrieval difficult in the CrewAI project? After explaining the basics of RAG, Jakub shared his experience from a project using CrewAI, a framework for building AI agent-based systems. He discussed three problems encountered in the configuration used: overly large text chunks, the absence of an additional relevance assessment and errors in PDF text extraction. 2.1 Overly large document chunks In the configuration Jakub described, text was split into chunks of 4,000 characters. The system retrieved five such chunks for each question, passing up to approximately 20,000 characters of source material to the model. A large chunk can contain information on several different topics, making it harder to match it to a specific question. The model generating the answer must then select the relevant information from the supplied content. In this project, the chunking approach therefore needed to be adapted to the structure of the documents and users’ questions. 2.2 No additional assessment of search result relevance Jakub pointed out that the configuration lacked reranking, which involves reassessing and reordering search results according to their usefulness for answering the user’s question. The system can first retrieve a larger number of passages, then assess them further to select those most useful for preparing an answer. Jakub presented this method as a potential improvement whose value should be evaluated by checking both answer quality and response time. 2.3 Incorrect text reading order in multi-column PDFs Another problem involved document text extraction. The tool read multi-column PDFs row by row, merging content from adjacent columns. This disrupted the order of sentences and made subsequent information retrieval more difficult. The resulting text was then split into chunks. The error therefore originated during data preparation and affected the subsequent stages of document processing. This example showed why evaluating RAG quality should begin with comparing the extracted text against the source document. 3. How does response time affect the choice between Classic RAG, Agentic RAG and Graph RAG? A key project requirement was a response time of around 4 to 5 seconds. Jakub discussed three RAG approaches in terms of data preparation costs, the ability to evaluate their operation and the time needed to handle a question. Approach How it works, as discussed during the session What to consider when choosing Classic RAG Retrieves passages from a knowledge base, optionally reranks them and passes the context to the model. Document chunking quality, retrieval relevance and the amount of context provided. Agentic RAG An agent selects tools and a retrieval method, running additional queries as needed. The ability to adapt retrieval to the question, along with the time and cost of additional operations. Graph RAG Retrieval uses a knowledge graph describing entities found in the source material and the relationships between them. The effort required to build and maintain the graph, and how useful the relationships are for answering users’ questions. In an agentic approach, the model can use several tools, such as vector search, keyword search or filtering by metadata describing the document. Additional steps allow the system to expand its search for information, while their number and sequence affect response time. In the graph-based approach presented, some of the work takes place when building the knowledge base. Entities and the relationships between them are extracted from the text. This mechanism also underpins GraphRAG as described by Microsoft. Jakub highlighted the costs of this preparation and the difficulty of manually analysing a complex graph. In this project, the response time requirement favoured further development of classic RAG. The team focused on document chunking and context selection. 4. How does hierarchical document chunking help preserve context? The solution organised the material into three connected levels: pages, paragraphs and sentences. The system retained information about which paragraph each sentence belonged to and which page contained that paragraph. Content was represented in the vector database at different levels of detail. This allowed retrieval to identify both individual sentences and larger passages containing the required information. According to Jakub, the additional cost of storing and processing these representations was acceptable given the volume of material in the project. Finding a relevant sentence made it possible to retrieve its entire paragraph and provide the model with broader context. The system could also retrieve the whole page when needed. Suppose a user asks about the deadline for appealing an exam result. The system finds a sentence specifying the deadline, then retrieves the entire paragraph explaining when the appeal period begins and how to submit an appeal. This allows the model to account for these conditions in its answer. 5. How can you evaluate RAG quality using your own data? In the final part of the presentation, Jakub emphasised the importance of a benchmark, a set of tests used to compare different versions of a solution. He discussed checking retrieval results against information labelled by a human and using a language model to evaluate answers. In practice, it is useful to assess two stages separately. The first concerns retrieval: did the system return a passage containing the required information? The second concerns the answer: did the model use the supplied material correctly? This distinction helps identify which stage needs improvement. In additional information shared after the session, Jakub clarified the testing method and results. The test set included questions covering the full scope of the documentation, along with real user questions collected anonymously during a prototype launch at the beginning of the year. Answer accuracy increased from around 70% to around 98%, an improvement of approximately 28 percentage points. This result applies to the internal test conducted in this project. According to Jakub, the solution also maintained a fast response time. When he shared these details, the chatbot had completed internal testing, and the company planned to make it available to a subset of customers. Reducing the support team’s workload and making information easier to access remained deployment goals. Assessing whether those goals have been achieved requires data from actual use. The embedding model is another component to evaluate. Jakub noted that its selection should take into account the language of the source material and its performance on the team’s own dataset. The choice of this model affects which passages the system retrieves before it begins generating an answer. 6. What can companies implementing a chatbot learn from this experience? The project shows how specific requirements guide RAG development. The expected response time helped narrow down the choice of solution, document analysis revealed problems with text extraction and chunking, and an internal test made it possible to assess the impact of the changes. When planning a similar implementation, it is worth addressing five areas: Source material: check whether document text extraction preserves meaning and reading order. Document chunking: adapt chunk size and the connections between chunks to the structure of the material. User questions: prepare a test set that reflects the tasks the chatbot is intended to support. Response time: establish expectations and account for them when comparing approaches. Quality assessment: check both the relevance of the retrieved information and how it is used in the answer. Let’s talk about AI in your company TTMS is home to experts who, like Jakub, combine technical knowledge with experience from client projects. During Tech Talks, they share solutions tested in practice and apply what they have learned to subsequent implementations. Are you planning a chatbot that uses your company’s documentation, or looking to improve the answer quality of an existing tool? Let’s talk. We will review your materials, users’ needs and business goals to recommend an appropriate way to use AI. Contact the TTMS team! What documents can a RAG chatbot use as a knowledge base? A RAG chatbot can use company policies, product manuals, procedures, FAQs and other materials containing information relevant to its users. Sources may include PDFs, Word documents, website content and knowledge base articles, depending on the integrations available. Scanned documents require optical character recognition (OCR) to turn images of text into searchable content. Tables, diagrams and complex layouts may need additional processing to preserve their meaning. Before adding documents, check that they are accurate, current and approved for the intended audience. Clearly structured materials help the system retrieve information and provide useful context for its answers. How do you keep a RAG chatbot’s knowledge base up to date? Keeping a RAG chatbot up to date requires a process for detecting and processing changes in its source materials. Depending on business needs, updates can run on a schedule or be triggered when a document is added, edited or removed. The system then updates the searchable content and its associated representations, such as embeddings. Version information and effective dates help distinguish current guidance from older material. Deleted or superseded documents should also be removed from active search results, and cached answers may need refreshing. Assigning an owner to each content area helps ensure that someone remains responsible for the information the chatbot uses. Can a RAG chatbot provide sources for its answers? Yes, a RAG chatbot can include links, document titles, page numbers or quoted passages alongside its answers. This requires the system to preserve source information when processing documents and connect retrieved passages to the response. Useful citations let users open the relevant material and check the context for themselves. The system should also verify that each citation supports the claim it accompanies. A source link alone provides no guarantee that an answer accurately reflects the document. During testing, teams should check both answer quality and citation accuracy, including whether users can access the referenced material. How can a RAG chatbot respect access permissions for company documents? A RAG chatbot can use the signed-in user’s identity and access rights to determine which documents it may retrieve. Permission checks should happen before restricted content reaches the language model. The same controls need to cover document previews, citations and any cached responses that contain protected information. When access rights change in a source system, those changes must also be reflected in the chatbot’s retrieval process. Teams should test the solution using accounts with different roles, including users with limited access. These checks help confirm that each person receives answers based on information they are authorised to view. What should a RAG chatbot do when it cannot find an answer? When the available documents provide insufficient information, a RAG chatbot should clearly explain that it cannot answer reliably from its sources. It can ask a clarifying question if the request is ambiguous or suggest a related document that may help. For questions requiring further assistance, it can direct the user to the appropriate team or support channel. The system needs explicit rules for handling incomplete, conflicting or missing information. Testing should include questions whose answers are absent from the knowledge base, so the team can assess this behaviour. Reviewing unanswered questions can also reveal gaps in company documentation and priorities for future updates.
ReadChatGPT for financial services: what does combining GPT with professional data sources offer?
On 10 September 2026, OpenAI announced ChatGPT for Financial Services, a solution combining GPT-6 Astra with professional financial data sources and tools for preparing analyses. The product was developed in collaboration with Morgan Stanley and Evercore. It is designed for financial institutions, with an initial focus on investment banking and equity research. It allows analyst teams to find data, perform calculations and prepare client materials in one place. This could reduce the time spent gathering information and transferring it between tools. In this article, you will learn: what data and features ChatGPT for Financial Services offers, what preparing a company analysis with GPT could look like, why metric calculations and data sources need to be checked, which stages require an analyst’s review, how to assess whether implementation is worthwhile for your company. How does ChatGPT for Financial Services support analysts? Materials published by OpenAI and its data providers describe several specific use cases: Comparing companies. Daloopa, a provider of company financial data, makes selected data and metrics available for comparing business performance. Source references help analysts verify where the figures come from. Finding companies that meet specific criteria. Daloopa also describes searching for companies by business activity or geographical region. The resulting list can provide a starting point for further market analysis. Preparing client materials. OpenAI describes creating valuation models, research notes and presentations using company templates for Excel, Word and PowerPoint. ChatGPT for Financial Services provides access to selected data from Daloopa, PitchBook, LSEG News and Crunchbase. OpenAI is also developing integrations that will allow institutions to use data covered by their existing subscriptions. These include S&P Capital IQ, LSEG, MSCI, Dow Jones Factiva and Moody’s. From data to company analysis: five steps in the workflow Let’s walk through preparing a comparison of two industrial companies ahead of a client meeting. The analyst needs to assess profitability, explain material differences and prepare a short note with a results table. Using fictional data, we will show what to check during a pilot, from selecting information to approving the final material. 1. Defining the question and scope of the comparison First, we establish which period to compare: the last full year, a six-month period or the trailing twelve months. We also check whether the figures cover the entire corporate group or an individual company, which currency they use and how each metric was calculated. In our example, we use consolidated data for both groups for the same calendar year. Amounts are stated in millions of euros. Before comparing results, we need to check the start and end dates of the reporting periods. One company’s financial year may end in December, while another’s ends in March. The documentation for the US SEC’s EDGAR database also highlights these differences. The analyst must then decide how to account for the mismatch and whether additional data is needed. The agreed approach should be recorded in the instructions for the AI and included with the completed analysis. This gives the model clear guidance and helps the reviewer understand which data was compared and why. 2. Gathering data and identifying its sources For each important figure, record the company and period it relates to, the units used and how it was calculated. A reference to the specific table or explanatory note in the report is also needed. Keeping the source document and its retrieval date makes it easier to review or update the analysis later. According to OpenAI’s description, ChatGPT for Financial Services lets users locate specific tables and document passages, highlighting the information used in the analysis. In our example, we compare EBITDA, or earnings before interest, taxes, depreciation and amortisation. The reviewer should be able to trace a reported value back to the company’s report and check how it was calculated. This also allows them to confirm that the figure covers the correct period and scope of operations. 3. Aligning definitions before comparing margins Companies may report adjusted EBITDA that excludes selected costs. These adjustments increase the value of the metric. Before comparing profitability, it is therefore necessary to check which adjustments have been applied. The US SEC also highlights differences in how individual companies calculate financial measures. Let’s look at two fictional companies. We assume that both calculate EBITDA before adjustments using the same principles. Company A then adds back EUR 4.59 million in costs that it excludes when calculating adjusted EBITDA. As a result, the metric rises from EUR 27.57 million to EUR 32.16 million. Company B has no such costs, so its figure remains unchanged. Illustrative example. Consolidated data for the same calendar year. Amounts are stated in EUR million. Item Company A Company B Revenue 229.73 183.78 EBITDA before adjustment 27.57 23.89 Costs excluded when calculating adjusted EBITDA 4.59 0.00 Adjusted EBITDA 32.16 23.89 EBITDA margin before adjustment 12% 13% Adjusted EBITDA margin 14% 13% After the adjustment, Company A’s margin is 14%, exceeding Company B’s margin of 13%. Before the adjustment, Company B has the higher margin: 13% compared with 12%. In this example, the treatment of costs determines which company has the higher EBITDA margin. The analyst should therefore check which costs make up the EUR 4.59 million adjustment and whether they also occurred in previous years. This helps them assess whether excluding these costs is justified for the analysis being prepared. They can also present both scenarios and explain the difference to the client. AI can help gather data and recalculate margins, while the expert assesses whether the adjustment is justified and how it affects the conclusions. 4. Verifying calculations in the spreadsheet In our example, simply divide EBITDA by revenue: 27.57 ÷ 229.73 gives a margin of approximately 12%, while 32.16 ÷ 229.73 gives approximately 14%. The displayed amounts are rounded; the spreadsheet should retain full precision for its calculations. More complex analyses may require currency conversion, alignment of reporting periods or the preparation of several forecast scenarios. The reviewer should be able to trace each of these steps. It is therefore worth asking AI to create a spreadsheet in which source data, assumptions and formulas are clearly separated. The analyst can then check the calculations and see how changing a single value affects the result. To test the spreadsheet, you can halve the adjustment, reducing it from approximately EUR 4.59 million to EUR 2.30 million. Company A’s adjusted EBITDA should then be approximately EUR 29.86 million, with a corresponding margin of 13%. These amounts are rounded for presentation; the spreadsheet should calculate the change using unrounded values. After making this change, check the comparison table and the commentary on the results as well. Both companies would now have the same margin, so the conclusion that Company A has a higher margin would need updating. This is a simple way to assess whether the calculations and accompanying text remain consistent. 5. Preparing client materials and reviewing conclusions The completed analysis can be presented in the company’s preferred format. According to OpenAI’s description, an administrator can share Excel, Word and PowerPoint templates with the team for the tool to use when creating documents and presentations. In our example, the client should receive a results table and a short explanation of how the cost adjustment affects the margin comparison. The expert reviewing the material checks whether the conclusions match the calculations and answer the client’s question. If anything needs clarification, they can request a further explanation or another version of the analysis. Time measurements should also include reviewing the material and making corrections before approval. How can you protect data and preserve a record of the analysis? When preparing a client analysis, the team may use public reports, paid databases and confidential documents. It is necessary to establish who can access this information, where it will be stored and who can receive the finished material. According to the ChatGPT Work security documentation, business data is encrypted and is not used to train models by default. Data retention periods, processing locations and the scope of recorded activity depend on the settings and connected services. Before implementation, check which user and tool actions are logged and which records can be exported. During the pilot, keep the source documents, successive versions of the spreadsheet and the final material, together with a record of who approved it and when. Then check whether this documentation allows you to reproduce the calculations and trace the approval of the analysis. Separately, verify whether system logs allow user and tool activity to be traced to the extent required by the company. How can you assess whether implementation is worthwhile? Start by choosing a task the team performs regularly, such as updating a company comparison after quarterly results are published. Before testing, measure how long this analysis takes using the existing method and define its quality requirements. These findings will provide a baseline for comparison with AI-assisted work. The time needed to review and correct the analysis must be added to the preparation time. OpenAI highlights this in its guidance on assessing the business value of AI, also recommending that implementation and ongoing usage costs be included. In practice, it is worth comparing: Metric What does it tell us? Time from starting the task to approving the analysis Does the client receive the finished material sooner? Include waiting time between stages. Total time spent by the analyst and reviewer Does the team spend fewer hours on preparation, review and corrections? Number of errors affecting the results or conclusions Does the analysis meet the same quality requirements as the existing approach? Accuracy and completeness of source references Can the origins of key figures and information be verified? Time needed to update the analysis How efficiently can new data be incorporated and the calculations and conclusions that depend on it be updated? Cost per approved analysis What is the cost of the finished material, including team time, the tool, data and the share of implementation and maintenance costs allocated to that analysis? The test should cover several tasks of varying difficulty. Define the assessment criteria before it begins. Someone performing the same analysis for a second time already knows the data and some of the answers, which may shorten the time needed. It is therefore worth using comparable tasks and varying the order in which participants work with AI and with the existing method. If AI saves time, check how the team used it. They may have prepared more analyses, responded to clients sooner or reduced overtime. The implementation assessment should show separately how the time saved was used and whether company spending decreased, and by how much. When is it worth starting a pilot? Consider a pilot if the team regularly gathers data from multiple sources and updates similar analyses. Choose a task that takes analysts a significant amount of time, such as comparing data from company reports. Assign a person to lead the pilot and experts to review the results. If the team only occasionally analyses a few annual reports, check whether tools already approved for use within the company are sufficient. Where data retrieval and calculations are already automated, identify a specific task that the new tool could improve. OpenAI makes the product available to financial institutions that meet its access requirements and directs interested companies to its sales team. Pricing, detailed terms and availability for a particular institution in Poland must be confirmed with the provider. Availability information. The implementation decision should be based on the pilot results: the quality of the analyses, the time needed to prepare and review them, and the total cost of the work. The test will also show whether the tool provides access to the data the team needs. Want to explore where AI could improve analysts’ work in your organisation? Talk to the TTMS team about choosing a task for a pilot, connecting the necessary data sources and assessing the results. How does ChatGPT for Financial Services differ from analysing reports in ChatGPT? ChatGPT for Financial Services provides access to selected professional financial data directly within the tool. It also supports references to specific tables and document passages, as well as the preparation of materials using company templates. When assessing its suitability for a team, check whether the available sources cover the companies, periods and metrics the team needs. Does ChatGPT for Financial Services require separate financial data subscriptions? Selected datasets are included in the product. These cover some of the information supplied by the providers named by OpenAI. The company is also developing integrations intended to let institutions use data covered by their existing subscriptions. Before purchasing, confirm which data is included in the offering and which requires additional access rights. Can ChatGPT for Financial Services be used to analyse companies listed on the Warsaw Stock Exchange? This depends on the availability of data for individual companies. Check whether the tool provides their financial statements, relevant metrics and historical data. The launch announcement alone does not confirm full coverage of the Warsaw Stock Exchange. The best way to assess the product’s suitability is to test it on several companies the team regularly analyses. What should you do if data from ChatGPT differs from the figures in a company’s report? Start by comparing the sources, reporting periods, units and definitions of the metrics. A discrepancy may arise, for example, from using standalone rather than consolidated financial data, or from including EBITDA adjustments. Also check whether the company has published an updated report. The analyst should explain the discrepancy and document which value they used and why.
ReadChatGPT, an Integrated LLM, an SLM or Automation? How to Choose the Right AI for Your Business Process
In 2025, 20% of enterprises in the European Union used artificial intelligence, compared with 13.5% a year earlier. In Poland, the figure was 8.4%. The most common application was analysing written language, used by 11.8% of the companies surveyed. The authors of a study published in 2026 in Organization Science described the uneven boundaries of AI capabilities as a “jagged technological frontier”: tasks of similar difficulty for humans can pose very different challenges for a model. The next step requires answering a more specific question: which form of AI fits a particular task? In an experiment involving 758 consultants, participants using GPT-4 completed 12.2% more tasks and finished them 25.1% faster on average when the tasks fell within the model’s capabilities. For a complex task beyond those capabilities, the probability of reaching the correct solution fell by 19 percentage points. The authors of the study, published in 2026 in Organization Science, called this uneven boundary of AI capabilities the “jagged technological frontier”. Effectiveness therefore depends on matching the technology to the task, data, risk and approach to verifying results. One process may be best served by an enterprise LLM assistant. Another may require an application connected to a CRM, a knowledge base and an access control system. A third may benefit from a small model running locally. Operations governed by explicit rules can be handled by code, rules engines or robotic process automation (RPA). Traditional machine learning is suitable for tasks such as forecasting and data classification. 1. LLMs and SLMs in Business: Choosing the Model, Integrations and Where Data Is Processed Terms such as “LLM”, “deployed LLM” and “closed SLM” combine several layers of technology. In a business context, it helps to separate four decisions: Way of working: does an employee interact with a ready-made assistant, or does the process start automatically? Scope of integration: does the solution work with materials supplied by the user, or does it retrieve data and perform actions in company systems? Model type: does the task require the broad capabilities of a large language model, or would a specialised SLM be sufficient? Processing location: does the model run as a cloud service, in a dedicated environment, in a private cloud, on premises or directly on a device? SLM stands for Small Language Model, which typically requires fewer computing resources. The term “closed system” needs clarification: it may refer to restricted access, an isolated environment or data processing within the organisation’s own infrastructure. A large language model can run in a private environment, while an SLM can be available through a public API. Model size alone does not determine how data is secured. 2. Six Ways to Use AI in Business Processes Approach How It Works Best Fit Key Metric Enterprise LLM assistant An employee assigns a task and checks the result in an approved environment, such as ChatGPT Business or Enterprise Analysis, drafting, summarising, developing alternatives and ad hoc work Time saved per task after accounting for review and corrections Integrated LLM or RAG application The model uses company sources, rules, permissions and integrations Repeatable processes involving documents, knowledge and data from business systems Cost per successfully resolved case AI agent The solution plans its next steps, selects tools and pursues a goal within a defined scope Multi-step processes involving exceptions and a dynamic sequence of actions Percentage of tasks completed correctly SLM A smaller model handles a narrow range of tasks in the cloud, on a company server, at the edge or on a device High task volumes, a fixed subject area, short response times and offline operation Quality compared with an LLM within a specified cost budget and p95 response time limit Private or on-premises deployment An LLM or SLM runs in a controlled environment, private cloud, company network or on a device Requirements relating to data residency, business continuity, connectivity, infrastructure or security policies Compliance with requirements, quality, availability and total cost of ownership (TCO) Rules, RPA or traditional ML The process is defined through code, conditions, a predictive model or a state machine Calculations, transactions, fixed workflows and unambiguous decisions Accuracy, repeatability and completeness of the audit trail In a mature deployment, these approaches often work together. The language model interprets a message, rules check the conditions, an application retrieves data, a person approves the action, and the transactional system records the result. 3. Which Tasks Are Suitable for ChatGPT or Another LLM Assistant? A ready-made LLM assistant supports tasks where an employee is responsible for checking and using the result. The user initiates the work, provides context, evaluates the response and decides how to use it. The output takes the form of a draft, recommendation, analysis or working document. Examples include: preparing a first draft of a report, message, presentation or article, summarising documents and correspondence, comparing several materials supplied by the user, developing questions, scenarios and alternative solutions, translating content for subsequent review, exploring data and explaining findings, organising meeting notes, drafting a procedure or action plan. This approach works well in processes where every response undergoes human review, the result can easily be corrected or withdrawn, and the task does not require automatic writes to a critical system. Wide variation in the source material and the need for language-related work further increase the usefulness of an LLM. Security depends on the approved product and its configuration. OpenAI states that data from ChatGPT Business, ChatGPT Enterprise and the API is not used to train models by default. The API data controls documentation also describes separate retention policies, including standard abuse monitoring logs and a Zero Data Retention option for eligible customers. Before using the solution, an organisation should review data classification, contractual terms, processing region, retention, administrator permissions and policies for connected applications. For more examples, see our overview of 15 ChatGPT integrations with business applications. 3.1 How Can You Measure Time Savings and the Quality of Work with an LLM? Relying solely on employee surveys can overstate the benefits. In a METR experiment, experienced developers took 19% longer to complete the tasks studied when using AI tools, yet afterwards still estimated that AI had made their work 20% faster. The study involved 16 participants and 246 real tasks in repositories they knew well, so its findings apply to that specific setting. The methodological lesson has broader relevance: actual task duration and quality need to be measured before deployment. For an LLM assistant, useful metrics include median task completion time, the proportion of outputs accepted without changes, average correction time, quality assessed against consistent criteria, frequency of use and the number of cases in which users return to their previous way of working. 4. When Should You Integrate an LLM with Company Systems, and When Should You Deploy an AI Agent? Integration becomes justified when the value of a process depends on current company data, a repeatable workflow and coordination across several systems. The model then receives controlled access to documents, a knowledge base, CRM, ERP, a ticketing system or email. The application verifies the user’s identity, controls which data is shared, defines the response format, checks results, logs actions and routes selected operations for approval. A typical integrated AI application consists of five layers: Input: a message, document, form, system event or record. Context: data retrieved in line with access permissions, often using RAG. Model: an LLM or SLM selected for the specific stage. Validation: rules, completeness checks, source verification and risk classification. Output: a response for a person, a draft record or an approved action in a system. Use cases requiring this architecture include finding answers in an internal knowledge base, reviewing contracts against a company’s risk checklist, preparing quotations using CRM data and a price list, classifying support tickets, onboarding employees, checking procurement documents and drafting responses based on customer history. RAG retrieves up-to-date passages from approved sources and adds them to the context used to generate a response. An AI agent represents a further level of integration. It receives a goal, selects tools and plans a sequence of actions. According to the current Google Cloud guidance on agentic architecture, agents are suited to open-ended, multi-step problems that require external data and a degree of autonomy. An application with a predefined sequence of steps is usually sufficient for a single summary or translation. We explore the role of an advanced model as a reasoning layer connected to tools, data and permissions in our article GPT-5.5 for Business: A New Era of AI Agents. The scope of an agent’s autonomous actions can be expanded when test results confirm the required effectiveness and safety. Its permissions should be limited to the functions needed for the process, read access should be separated from write access, and actions with significant consequences for the company or customer should require approval. OWASP identifies excessive functionality, excessive permissions and excessive autonomy as the three main causes of Excessive Agency risk. The principle of least privilege limits the consequences of misinterpretation, fabricated information or an attack that feeds malicious instructions to the model (prompt injection). 5. Which Business Processes Are Suitable for a Small Language Model (SLM)? A Small Language Model uses fewer parameters and computing resources than a large language model. According to Microsoft Azure, this can support faster responses, lower infrastructure requirements and data processing close to where the data originates (edge computing), for example on industrial devices. A specialised SLM can be suitable for tasks with a fixed subject area, predictable inputs and a clearly defined output. An SLM is worth testing when a process meets several of the following conditions: a fixed set of categories, intents, fields or response types, a high and predictable volume of requests, a requirement for a short response time measured as p95, the threshold within which 95% of responses are completed, limited hardware resources, a need to operate offline or directly on a device, access to data from the relevant business domain and a set of reference answers, the ability to escalate difficult cases to a larger model or a person. Examples include classifying support tickets into 30 fixed categories, identifying a user’s intent, extracting field values from a single document type, generating short responses within a tightly defined subject area or analysing messages locally on an industrial device. The choice depends on testing with company data. An SLM should meet the required targets for quality, response time and cost per correctly handled case. For example, a company might require the smaller model to retain at least 98% of the reference LLM’s quality score, reduce the cost per correct result by at least 20% and stay within the p95 response time limit. These are illustrative decision thresholds that the process owner sets before the pilot. 5.1 When Should You Deploy an LLM or SLM On Premises or in a Private Cloud? A private or on-premises deployment may be driven by requirements relating to data sovereignty, security policies, business continuity, network latency or operation without internet access. Such a deployment can use an SLM or a larger model. An enterprise application can also use a managed API with encryption, retention controls, an appropriate processing region and contractual provisions governing data handling. A cascading architecture can be the most effective approach. Microsoft describes a hybrid model in which an SLM handles routine queries and passes more complex cases to an LLM. In a business setting, it is worth adding a third route to this cascade: referring the case to an employee when the result is uncertain, the risk is high or required data is missing. 6. Which Processes Should You Automate with Rules, RPA or Machine Learning? Processes governed by fixed rules require a clearly defined sequence of actions and conditions for carrying them out. AWS documentation on orchestration distinguishes between rule-based workflows, where successive states and transitions are explicitly defined, and agentic orchestration, where a model interprets the goal and dynamically selects tools. Both layers can operate within a single application. Process Recommended Mechanism Role of the Language Model Calculating tax, pay or a discount Code and a rules engine Explaining the result or interpreting the user’s query Executing a payment, refund or limit change A transactional process with access controls Identifying intent and preparing data for approval Granting or revoking permissions Identity and access management (IAM), role-based rules and approvals Handling a request expressed in natural language Checking that all required fields are complete A schema validator Extracting fields from an unstructured document Predicting customer churn or forecasting demand Traditional machine learning Explaining contributing factors and drafting communications Interpreting a free-form message An LLM or SLM Classifying intent and passing data to a controlled workflow In a financial process, an LLM can read a message, identify the request and prepare a proposal. Rules check the balance, limits, customer status and required approvals. The transactional system executes the operation once the conditions are met. Each layer performs a task for which clear criteria for correctness can be defined. 7. How Do You Match AI to a Business Process? Four Assessment Criteria An initial assessment can be carried out during a short workshop. The process owner evaluates the process across four areas, assigning a score from 0 to 3 in each. The individual scores help define requirements for the model, integrations, safeguards and infrastructure. Dimension 0 1 2 3 Complexity of content interpretation Fixed fields and rules A fixed set of categories Interpreting context Combining information and reasoning across multiple sources Integration and autonomy No access to systems Reading from a single source Reading from multiple sources or preparing data to be written to a system Transactions and dynamic tool selection Impact of an error Easily reversible Limited operational cost Significant financial, legal or reputational impact Critical or irreversible consequences, or an impact on rights and safety Infrastructure and data requirements A managed cloud meets the requirements A specific region or retention policy is required A private network or strict latency limit Operation offline, in an air-gapped environment, at the edge or directly on a device The scores can be interpreted as follows: Content interpretation complexity of 0-1 with stable rules: code, workflows, RPA or traditional ML. Complexity of 2-3, integration of 0-1 and error impact of 0-1: an enterprise LLM assistant with user review. Complexity of 2-3 and integration of 2-3: an integrated LLM application, RAG or an agent. Complexity of 1-2, a narrow domain and infrastructure and data requirements of 2-3: an SLM as a candidate for comparative testing. Error impact of 2-3: approval by an authorised person, safeguards based on predefined rules and a complete activity log, regardless of model type. The table helps identify solutions for a pilot. The final decision follows a comparison of their performance on the same set of real cases. 8. ChatGPT, Integrated LLMs, SLMs and Automation: Business Use Cases Example Process Recommended Architecture Key Performance Indicator Human Oversight Drafting marketing content An enterprise LLM assistant Median time saved and percentage of outputs accepted Approval of every publication Summarising a meeting and listing action items An enterprise assistant with access to an approved source Completeness of action items and number of corrections Verification of task owners and deadlines Answering questions about internal procedures An integrated LLM with RAG and source citations Percentage of answers grounded in sources and accuracy of citations Escalation when no source is available Assigning support tickets to predefined queues An SLM or classifier, with an LLM for exceptions Macro-F1 and the proportion of priority tickets correctly identified Review of uncertain cases Reviewing contracts against a company’s risk checklist An integrated LLM with RAG, rules and logging Rates of detected and missed risky clauses Decision by a lawyer Extracting fields from a single invoice type OCR, traditional ML or an SLM, combined with rule-based validation Field-level accuracy and cost per document processed Review of exceptions Preparing a quotation using CRM data and a price list An integrated LLM, RAG and price retrieval governed by predefined rules Preparation time and percentage of quotations requiring commercial corrections Approval of pricing and terms Checking refund eligibility Process rules Compliance with policy and time to decision Handling exceptions Executing a refund A transactional workflow with authorisation 100% accounting accuracy and a complete audit trail Depends on the amount and risk Analysing machine messages locally An SLM or specialised model running at the edge p95 response time, alarm detection rate and availability without an internet connection Escalation of critical alarms Revoking access for a departing employee IAM and a deterministic workflow Completeness of access revocation Approval in line with policy Handling a customer case across multiple steps An AI agent with a restricted set of tools Task success rate, correctness of tool use and percentage of cases referred to an employee Approval checkpoints for high-impact actions 9. How Do You Measure the Results of an AI Deployment? Quality, Time and Cost Metrics Measurement starts with the existing process. The Generative AI at Work study, involving 5,179 customer support employees, found an average increase of 14% in the number of issues resolved per hour. Less experienced employees saw the greatest improvement. Productivity defined this way has a clear numerator, denominator and comparison group. An enterprise pilot requires similar precision. Area Metric How to Measure It Scale Case volume Number of cases per month, seasonality and peak demand periods Time Case handling time Mean, median and p90 before and after deployment Quality Task success rate Percentage of cases meeting all criteria for correct task completion Usefulness Percentage of outputs accepted without substantive changes Proportion of outputs accepted without changes to their substance Oversight Percentage of AI decisions changed by an employee Proportion of decisions or proposals modified by an employee Risk Critical error rate Number of critical errors per 1,000 or 10,000 cases Classification Precision, recall and F1 score Measured separately for each important class, particularly rare events RAG Consistency of responses with sources Percentage of claims supported by a cited, up-to-date source Agent Correctness of tool selection and supplied parameters Whether the correct tool was selected and valid parameters were supplied Automation Percentage of cases handled entirely automatically Proportion of cases completed without manual intervention Performance Time from task initiation to the final result p50 and p95 of end-to-end task completion time Economics Cost per successfully completed task Total cost divided by the number of correct outcomes Stability Changes in system quality over time Variation in quality by time period, language, category and user type Google Cloud identifies cost per successfully completed task as a key metric for AI agents operating in real business processes. A model that costs USD 0.10 per run and achieves a 50% success rate incurs a model invocation cost alone of USD 0.20 per correct result. Human review, retries, integrations, monitoring and the cost of errors must also be included. 10. How Do You Calculate the ROI of an LLM or SLM Deployment? The full cost of the solution should include development and integration, the model or API, infrastructure, monitoring, updates, human review, corrections and expected losses resulting from errors. Cost per successfully completed task: C_success = (development cost allocated to the period + model/API + infrastructure + monitoring + human review + corrections + expected losses from errors) / number of successfully completed tasks Annual benefit: Benefit = value of working time saved + avoided correction and error costs + additional margin + avoided SLA penalties ROI: ROI = (Benefit - TCO of the AI solution) / TCO of the AI solution × 100% Suppose a team classifies 20,000 support tickets per month. Each ticket takes an average of 4 minutes, and the fully loaded hourly labour cost is PLN 120. The monthly cost of manual classification is approximately PLN 160,000. During the pilot, 88% of the system’s outputs are accepted without correction. Reviewing each output takes an average of 45 seconds, correcting the remaining cases takes 3 minutes each, and the monthly costs of the model, infrastructure, maintenance and amortised implementation total PLN 51,000. reviewing all cases: approximately PLN 30,000, correcting 12% of cases: approximately PLN 14,400, model, infrastructure, maintenance and implementation: PLN 51,000, total process cost after deployment: approximately PLN 95,400, monthly cost reduction: approximately PLN 64,600, or 40.4%. This example illustrates the calculation method using assumed values. A financial assessment of an actual deployment must account for changes in the cost of errors, seasonality, downtime, exception handling costs and the pace of employee adoption. For a revenue-generating process, the calculation should also include changes in margin, conversion rate or customer retention. 11. How Do You Run a Measurable LLM or SLM Pilot? Establish a baseline. Measure volume, time, quality, errors, escalations and the cost of the current approach over at least one full business cycle. Prepare a test dataset. Include typical tasks, difficult and rare situations, boundary cases and deliberate attempts to mislead the system. Google Cloud recommends a custom dataset that reflects the full range of intended uses. Set thresholds before testing. Document the minimum quality, maximum cost, acceptable latency, critical error limit and escalation rules. Compare several approaches. Include the current process, a capable LLM, a smaller model and a hybrid architecture. Run the system in shadow mode. AI generates outputs alongside the existing process. Decisions and actions continue under the established rules. This helps identify errors before AI is allowed to handle live operations. Deploy the system within a limited scope. Start with proposals and approvals. Expand automated actions based on the results. Monitor performance quality. Repeat testing whenever the model, AI instructions, tools, data sources, rules or types of input material change. 11.1 How Many Test Cases Do You Need to Evaluate an LLM or SLM? For a metric expressed as a proportion, a conservative sample size at a 95% confidence level and a margin of error of +/-5 percentage points is approximately 385 independent cases. A margin of +/-3 percentage points requires approximately 1,068 cases. The representative sample should be supplemented with a separate set of critical and edge cases. For rare errors, the so-called rule of three is useful. If no critical errors occur in 300 tests, the approximate upper bound on their true probability at a 95% confidence level is still around 1%. Zero errors in 3,000 tests reduces that bound to approximately 0.1%. High-risk processes therefore require much larger datasets and tests targeting specific threats. 11.2 When Should You Complete an AI Pilot and Move into Production? Example Criteria Process Example Criteria for Moving into Production Support ticket classification Macro-F1 of at least 0.90; recall for priority tickets of at least 0.99; override rate no higher than 8%; p95 response time no longer than 2 seconds Internal knowledge base Acceptance rate of at least 85%; citation accuracy of at least 98%; no unsupported claims in the critical test set; p95 response time no longer than 8 seconds Draft quotation Median preparation time reduced by at least 30%; at least 75% of drafts accepted with minor changes; 100% of prices retrieved from an authorised source; human approval for every quotation These values illustrate how to define acceptance criteria. The process owner sets the thresholds according to the cost of errors, required quality and the organisation’s risk tolerance. 12. How Do You Match Human Oversight to AI Risk? NIST defines risk as a combination of the likelihood of an event and the scale of its consequences. This principle translates general concerns about AI into a measurable model: error frequency, the value of funds or resources at risk, the ability to reverse an action, the time needed to detect an error and the cost of correcting it. In the consolidated text of the EU AI Act, requirements for high-risk systems include continuous risk management, appropriate levels of accuracy, robustness and cybersecurity, and effective human oversight. Oversight measures should be proportionate to the risk, degree of autonomy and context of use. Testing should use predefined metrics and thresholds appropriate to the intended purpose. In business practice, this means assigning a specific person responsibility for approvals, monitoring and intervention. The interface should display sources, actions taken and the level of uncertainty, and the user must be able to stop the process. The risk of serious consequences from an error justifies restricting model permissions, adding further checks and extending shadow-mode testing. The regulatory classification should be assessed separately for the intended use and the organisation’s role. 13. How Can You Combine LLMs, SLMs and Rules in a Single Business Process? Combining several technologies allows each stage of a process to be handled appropriately. For example, an SLM identifies the topic of a customer’s request, while an LLM drafts a response using the contact history and documents retrieved through RAG. If the case involves a refund, the system checks conditions and limits against established rules, then routes the operation for execution or approval by an authorised employee. This approach combines automation with oversight of decisions that have financial consequences. At TTMS, we can help you assess where a similar solution would deliver the greatest benefit. We will start with the tasks that take up the most time: repetitive activities, searching for information or correcting errors. During a consultation, we will examine your workflows, available data and existing systems. Based on this assessment, we will recommend the technology and pilot scope, then work with you to define the expected outcomes and how to measure quality, time and costs. Ready to choose your first process to improve? Book a consultation with us about implementing AI. Sources Eurostat, 20% of EU enterprises use AI technologies, 11 December 2025. Fabrizio Dell’Acqua et al., Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality, Organization Science. Erik Brynjolfsson, Danielle Li, Lindsey Raymond, Generative AI at Work, NBER Working Paper 31161, 2023. Joel Becker, Nate Rush, Beth Barnes, David Rein, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, METR, 10 July 2025. Microsoft Azure, What Are Small Language Models (SLMs)? Microsoft Azure, Boost processing performance by combining AI models, 8 January 2025. OpenAI, Enterprise privacy at OpenAI. Google Cloud, The KPIs that actually matter for production AI agents, 26 February 2026. NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, July 2024. European Union, Regulation (EU) 2024/1689, consolidated text of 27 July 2026. OWASP GenAI Security Project, LLM06:2025 Excessive Agency, 2025. FAQ Do You Need to Run Your Own AI Model to Work with Company Data? Running your own model is one of several options. Companies can use an approved business environment, a managed API, a private cloud, dedicated infrastructure or a model running locally. The choice depends on data classification, processing location, retention, encryption, user identities, industry requirements and the agreement with the provider. For example, OpenAI states that business customer data is not used for training by default and offers additional retention controls for eligible API customers. Organisations should document the data flow for their specific configuration, as the model’s name does not describe the full security architecture. Can RAG Replace Fine-Tuning a Model on Company Data? RAG and fine-tuning address different needs. RAG retrieves current information from a controlled source when generating a response, making it well suited to knowledge bases, procedures, documentation and frequently updated content. Fine-tuning uses examples to adapt a model’s behaviour to a particular style, format or specialised task. AWS documentation comparing RAG and fine-tuning recommends starting with RAG for a question-answering system based on your own documents, particularly when up-to-date information and source references matter. Both approaches can work together when a process requires current knowledge and consistent behaviour within a specific domain. Can a Single Process Use Both an LLM and an SLM? Yes. A routing component can direct routine, clearly identified cases to an SLM and complex cases to a larger LLM. A third route passes cases to a person when the system detects missing data, low confidence or a high level of risk. Another option is to divide the process by function: an SLM classifies the document, an LLM prepares an explanation, code calculates values, and a workflow records the approved decision. This setup helps control costs and response times while retaining access to more advanced capabilities for difficult cases. How Many Examples Do You Need for an LLM or SLM Pilot? The number depends on the required measurement precision and how rare the errors are. When measuring the proportion of successful responses, a sample of approximately 385 independent cases provides an approximate margin of error of +/-5 percentage points at a 95% confidence level under conservative assumptions. A margin of +/-3 percentage points requires approximately 1,068 cases. The random sample should reflect the actual distribution of cases across languages, channels and user types, in proportion to their volumes. A separate test set should cover critical, edge and rare cases, along with attempts to manipulate the system. How Often Should You Retest an Application Built on a Language Model? A full evaluation should be run whenever the model, prompt version, tool, permissions, data source, business rules or input format changes. Production use also requires continuous monitoring of key performance indicators and regular regression testing. The frequency depends on the level of risk and how quickly the process changes. An application handling marketing content may follow a different schedule from a system supporting financial decisions. A practical approach is to run automated tests after every technical change, review trends monthly and conduct a business evaluation quarterly, with shorter cycles for high-risk applications.
ReadGPT-6 Astra in Microsoft 365 Copilot: Access, Tasks and Cowork Costs
Does your company use Copilot, and would you like to try GPT-6 Astra? OpenAI’s model is also available in Copilot Cowork. This means you can try it when working with documents, email and calendars in Microsoft’s environment. Access to Astra depends on your organisation’s licences and settings, while tasks performed in Cowork are billed based on credit usage. What does your administrator need to enable? Which tasks can you delegate to Astra in Cowork, and how are they handled in ChatGPT Work? Below, we explain access requirements, differences in working with files and billing rules. For guidance on choosing an assistant for your organisation, see our comparison of Microsoft Copilot and ChatGPT for business. 1. What Does GPT-6 Astra Bring to Copilot Cowork? Microsoft lists GPT-6 Astra among the models available in Copilot Cowork. Users select a model from the list enabled by their organisation. The default Auto setting lets Cowork choose a model for the task; a label next to the response shows which model was used. Selecting Astra applies to work within Cowork. The availability of a particular model in other Copilot features needs to be checked separately. GPT-6 Astra is another model you can assign tasks to in Cowork. Cowork itself provides the tools for finding information, creating files and taking action in Microsoft 365. Your choice of model may affect how information is analysed, the level of detail in the response and the time taken to complete the task. When evaluating Astra, check whether it handles an existing task better: whether it brings together findings more accurately, accounts for exceptions and produces a result that requires fewer revisions. Work IQ gives Cowork access to the context of your organisation’s work. When preparing a project summary, the information needed may be spread across documents, correspondence and meeting materials. Cowork can search for the organisational resources required for the task. Before trying it, check that the employee’s account has access to the relevant materials and that they include the latest decisions and updates. This determines which information Astra will use to produce its result. We discuss the model’s test results and examples of its use in our article GPT-6 Astra: Impressive Achievements and New Possibilities for Business. 2. How Can You Access Astra in Copilot Cowork? For business users, Microsoft describes Cowork as a service that requires a Microsoft 365 Copilot licence and usage-based billing for task execution. An administrator then needs to configure employee access. There are two separate settings to configure: Access to Cowork. The employee must belong to a group covered by a spending policy that includes Cowork. The administrator configures this in the Microsoft 365 admin centre under Copilot, Cost Management, Configuration. This is where they specify the users, budget and billing method. Access to models provided by OpenAI. In the Copilot settings, the administrator specifies which users can use OpenAI as a Microsoft subprocessor. Once you have access, open Cowork and select Astra from the model list. For your first task, check the model label next to the response. If an employee can see Cowork but cannot find Astra, the administrator should check the model provider settings. To make Cowork available to a specific team, the administrator must grant access to the relevant user group. A low credit limit restricts spending while still allowing employees covered by it to get started. 3. Astra in Cowork and ChatGPT Work: Differences in Task Execution Preparing a report involves finding up-to-date data, processing it and saving the result somewhere the team can access. At each stage, the tools available to the model matter. The comparison below shows how the two environments work with the materials needed for a task. Working with Materials in Copilot Cowork and ChatGPT Work Task Component Copilot Cowork ChatGPT Work Finding materials Searches Microsoft 365 resources accessible to the user, including email and files. Plugins can provide access to additional sources. Uses files provided for the task and information retrieved through enabled apps and authorised accounts. Working on documents Creates and modifies documents, spreadsheets and presentations. Output files are saved to the workspace in OneDrive or SharePoint. Creates and edits files. Transferring them to another system depends on the operations supported by the connection to that system. Files stored on the computer A file can be uploaded to the session. Cowork does not edit files directly on the user’s drive. Work in a supported desktop app can use local files once the appropriate access has been granted. Using an application through a browser The local Edge browser uses the employee’s existing sign-in. The feature must be enabled by an administrator. Access depends on the browser tool selected and the permissions granted. A cloud task requires separate authorisation to access company resources. When working through a browser, you need to consider where the task is running. Cowork supports the local browser when its web version is open in Edge. This feature is currently unavailable in the Copilot desktop app and on mobile devices. Cowork and Edge must also use the same work account. If the computer goes to sleep, actions requiring the local browser may be paused. In ChatGPT Work, the model can use shared files and applications on the computer during a local task. A task launched in the cloud runs in a separate environment. If the required materials are stored only on the employee’s drive or are accessible through a company VPN, they need to be made available to that environment through a supported method. As it works, Cowork displays the successive stages of the task. You can interrupt the session, clarify your instructions or provide missing information. Before taking significant actions, such as sending a message or scheduling a meeting, Cowork asks for approval. The additional confirmations it requests also depend on permissions granted earlier. For your first trial, choose a task for which you can clearly identify both the source materials and where the result should be saved. You can find examples of responsibilities in sales, HR, finance and other departments in our overview of 10 practical uses of Microsoft Copilot in an organisation. 4. How Much Does It Cost to Use Astra in Cowork and ChatGPT Work? Your budget needs to cover both the subscription and the use of tools to carry out tasks. For a company that already has the appropriate licences, enabling Cowork primarily means budgeting for usage charges. Below are the public prices for selected business plans. Subscription prices. Charges for task execution are explained below. Plan Monthly Price per User Terms Microsoft 365 Copilot Business EUR 18.20; currently EUR 15.60 under a promotional offer Billed annually, excluding tax. A separate qualifying Microsoft 365 licence is required. Available for up to 300 users. Microsoft 365 Copilot for enterprise EUR 26 Billed annually, excluding tax. A separate qualifying Microsoft 365 licence is required. ChatGPT Business USD 20 when billed annually or USD 25 when billed monthly A minimum of two users. Public pricing in USD; the final amount depends on factors including taxes and the market where the subscription is purchased. ChatGPT Enterprise Custom pricing Usage limits and billing are defined in the agreement. The Copilot Business promotion applies to the first year with an annual commitment and runs from 1 July to 31 December 2026. 4.1 How Are Cowork Tasks Billed? Cowork charges for factors including model usage, context retrieval, tool calls and runtime. Usage is converted into Copilot Credits; under the published pay-as-you-go pricing, one credit costs USD 0.01. One thousand credits therefore cost USD 10. The cost of an individual task depends on the number of credits consumed. The selected reasoning level also affects usage. Cowork offers Light, Medium, High, Extra High and Max settings. A higher level may increase task duration and credit consumption. For recurring work, check whether increasing this setting improves the result enough to justify the cost. 4.2 What Does Credit Usage Mean in ChatGPT Work? ChatGPT Work follows the usage limits and billing rules of the relevant plan. Under agreements based on a shared credit pool, tasks reduce the available balance. Credits already paid for under the agreement are covered by that payment. Additional charges may arise once those credits run out, if the agreement and settings allow work to continue. When comparing costs, use the same set of tasks and output requirements. Record usage, the number of retries and the extent of any revisions needed. Calculating the cost per successfully completed task shows how much you pay for a result your team can use. First, convert each service’s credit usage into a monetary amount using its own pricing. 5. What Data Protection Rules Apply to Astra in Cowork? In Copilot Cowork, Astra is provided by OpenAI as a Microsoft subprocessor. According to the documentation, this use of the model is governed by Microsoft’s terms and Data Protection Addendum, subject to specified exclusions. These services fall within the EU Data Boundary, with documented exceptions. Microsoft currently excludes them from its commitments to process data in a specific country. This detail matters to organisations that require processing exclusively in Poland, for example. When enabling Astra, the administrator should therefore consider the model provider’s policies and access to the materials used in the task. In ChatGPT Work, whether the task runs locally or in the cloud also matters. During a local task, file excerpts, screenshots and tool outputs may be sent to OpenAI. Company AI policies should account for this method of sharing information as well. 6. Which Task Should You Start with When Trying Astra? Start with a responsibility that regularly involves an employee gathering information and preparing material for other people. This workflow lets you assess both Astra’s analysis and the tools available in Cowork or Work. A weekly project summary is one example. A sample prompt for your own trial: Using the project folder [link] and correspondence about this project from the past seven days, prepare a report for the manager. List revised deadlines, pending decisions and the people responsible for next steps. Provide a source and date for each finding. If the materials contain conflicting information, show the discrepancy and explain what you need to resolve it. Save the report as a DOCX file using the attached template in the folder [link]. Draft a message to [recipients] with a link to the report. Leave sending it subject to my approval. Check whether the report reflects the latest decisions and updates, provides sources and dates, identifies conflicting information and assigns responsibilities correctly. Also assess whether it follows the template and whether the file has been saved in a folder accessible to its recipients. After a successful trial, you can consider running the task regularly. Cowork supports scheduled tasks and tasks triggered by events such as an email or a Teams post. By default, event-triggered tasks prepare actions for approval. We discuss how to design the entire process in our guide to business process automation with Copilot. 7. Prepare Your First Astra Tasks with TTMS Through our AI consulting services, we help you determine which data a task requires, which tools need to be made available and how to assess the result. We also analyse the required licences and usage billing arrangements. We combine consulting with AI solution design and the integration of business systems. TTMS was the first company in Poland to obtain accredited ISO/IEC 42001 certification for its artificial intelligence management system. The TÜV Nord Poland audit covered AI design and usage policies, including risk management and project documentation. Tell us which task you would like to delegate to Astra and which applications your team uses. Talk to TTMS about AI consulting for your business. GPT-6 Astra in Copilot Cowork: Frequently Asked Questions Does selecting Astra in Cowork change the model across all Copilot applications? The selection applies to work within Cowork. Microsoft describes a separate model selection option for this environment. To find out which model powers a particular feature in Word, Excel or Teams, check that feature’s documentation. When reviewing a Cowork task, you can see which model was used by checking the label next to the response. Will the same model give an identical response in Copilot and ChatGPT? The result may differ. The model works with the information provided by each product and uses its tools, instructions and reasoning settings. When comparing results, check which materials the model received and which actions it could perform. Only then can you meaningfully assess the differences in the outputs. Why can I see Cowork but cannot select GPT-6 Astra? Access to Cowork and access to OpenAI models are controlled by separate settings. Your administrator should confirm that your account is allowed to use models provided by OpenAI as a Microsoft subprocessor. The model list displayed in Cowork reflects the access granted by your organisation. Does a Microsoft 365 Copilot subscription cover all Cowork tasks? Cowork tasks incur additional usage-based charges. Copilot Credit consumption depends on factors including the model, information retrieval and tools used. Administrators can set spending policies for users and groups. Your budget should account for both the subscription and expected Cowork usage. Can Astra in Cowork edit a document saved on my computer? You can upload a document to a Cowork session. According to the current FAQ, the service does not open or edit files directly on your local drive. Cowork works with the materials you provide and files available in OneDrive and SharePoint. Support for the Edge browser is a separate feature. Will the same Astra model produce the same result in Cowork and ChatGPT Work? The result also depends on the available data, instructions, tools and reasoning settings. A task performed using the same model may therefore proceed differently in the two environments. Comparing results using the same materials will reveal differences in the completeness of the output and the actions performed.
ReadGPT-6 Astra: Impressive Achievements, New Opportunities for Business
OpenAI unveiled GPT-6 Astra on 3 September 2026, and early research findings and user reports show why the release is generating so much excitement. The new model solved mathematical problems that earlier GPT models and competing models failed to solve in the same test. Users have already tested GPT-6 across a range of tasks, from rebuilding a sales workflow in a CRM system and creating a detailed steam locomotive model to analysing a novel spanning more than 500 pages. What exactly has it achieved, and how can businesses put these capabilities to use? 1. GPT-6 Astra’s early achievements are impressive Early tests and user reports show how Astra handles complex tasks. These include findings from a test conducted by Epoch AI and accounts from people who put the model to work on their own projects. 1.1 Solving two previously unsolved mathematical problems Epoch AI tasked five models with solving 68 Erdős problems, giving each the same time and budget limits. Only the pre-release version of Astra solved two of them and produced solutions in a form that passed automated mathematical verification. GPT-5.6 Sol, GPT-5.5, Claude Fable 5.1 and Claude Fable 5 did not solve any of the problems in this test. Astra had already recorded other mathematical achievements. Before its release, OpenAI presented ten further results involving problems in mathematics and theoretical computer science. We covered them in our article “Astra, the future GPT-6: a new model from OpenAI?”. 1.2 A 3D steam locomotive model with 3,295 editable components Tom Krcha, a designer and creator of the AI-assisted interface design tool Pencil, gave Astra an old drawing of a steam locomotive. The task was to recreate the machine in Blender, an application for building 3D models, animations and scenes. Within a few minutes, it produced a model containing 3,295 separate objects. The wheels, axles, boiler components and other parts can be selected, moved and modified independently. Astra had to interpret a flat drawing, reconstruct the machine’s three-dimensional structure and preserve the relationships between thousands of components. According to Krcha, it did this mainly by writing Python scripts that built the geometry piece by piece. The resulting model can serve as a starting point for an animation or a game project. 1.3 Rebuilding a sales workflow directly in a CRM system Claire Vo, creator of the ChatPRD platform, gave Astra access to her customer relationship management system, or CRM. Working through Codex, the model was tasked with changing how new sales leads were handled. It had to understand the existing workflow, find the relevant settings and rebuild the rules in a visual editor. After the changes, the system automatically routed leads to Claire or Zach, inserted a link to book a meeting with the appropriate person and sent the draft message to Slack for approval. The new rules would also apply to future leads. 1.4 Adding new capabilities to a Bluetooth speaker with Astra Claire Vo also used the model to experiment with a small Divoom speaker fitted with a colour pixel display. She wanted to show her own images and messages on it. According to her account, the device had no public API, an interface that would allow other software to control it. This meant working out how to communicate with the hardware. Astra built an application that displayed drawings made with a computer mouse on the speaker’s screen. It then created a tool for controlling the speaker. The model also looked up information about the latest podcast episode and sent scrolling text and an animated graphic to the display. Vo noted that earlier attempts with other models had only allowed her to display a simple greeting. This time, she received custom software for controlling the device that she could develop further to suit her ideas. 1.5 Checking plot consistency in a novel spanning more than 500 pages Jakub Szczęsny of Antyweb gave Astra an extensive draft of his own book. The model was asked to check the chronology of events, the logic of the plot and storylines introduced in one part of the text and developed many chapters later. Analysing a manuscript of this length requires tracking the characters’ stories, the sequence of events, their motivations and the consequences of earlier decisions at the same time. Astra mapped these connections and flagged passages that needed further work, consistently checking the entire text for the specified issues. Szczęsny particularly valued its ability to connect information scattered across hundreds of pages. A similar skill is useful when reviewing contracts, project documentation and reports, where details in one section affect how others should be interpreted. 1.6 An AI agent completes the entire game Portal A creator publishing as CozyBlaze connected Astra to Portal, a spatial puzzle game in which players create passages between distant locations and use the laws of physics to overcome obstacles. The model received screenshots and information about the player character’s position and the direction they were facing. It used this information to plan moves, execute them and check the results. After approximately 23 hours and 43 minutes, including waiting time and technical interruptions, it reached the end credits. The creator developed custom controls and settings to make precise movement easier, and the game paused while the model was processing its next actions. The creator also resumed the session following service availability issues, while the agent made the gameplay decisions. Completing the game required interpreting the situation on screen, spatial awareness and hours of planning moves and checking their effects. These examples help explain the enthusiasm around Astra. The model analyses a problem, carries out successive actions and uses the results to guide what it does next. For businesses, this opens up the possibility of assigning AI more complex tasks involving information and applications. Comparisons with GPT-5.6 show where performance has improved and what those gains could mean for businesses. 2. GPT-6 Astra in business: what does better performance mean for companies? Businesses using AI also bear the cost of checking responses, correcting errors and stepping in when the model cannot finish a task. Better performance from the next generation could therefore make more tasks cost-effective to delegate to AI. Astra’s results give businesses reason to reconsider how responsibilities are divided, how existing systems are used and how much time employees spend reviewing AI output. GPT-6 Astra vs GPT-5.6 Sol: test results and their business implications Skill tested GPT-6 Astra GPT-5.6 Sol Business implications Successfully completing a task across several applications, AutomationBench 41.4% 28.77% Reason to test whether AI can handle a larger part of a process independently. Locating the correct elements on screen, ScreenSpot-Pro 92.7% 76.9% Greater precision when AI interacts with business software. Detecting bugs that require analysing several files, the more challenging CodeRabbit subset 57.1% 47.6% Better performance when analysing complex dependencies in software. The results come from different tests and measure distinct skills. AutomationBench: Zapier’s leaderboard as of 9 September 2026, with both models set to Max. ScreenSpot-Pro: OpenAI’s comparison. CodeRabbit: the more challenging subset of code reviews. The business implications are interpretations of the results; any savings need to be assessed within the company’s own process. 2.1 A broader range of tasks to delegate The more stages a task involves, the more opportunities there are for mistakes. The model needs to find information, apply the right rules, perform actions in the correct order and use the same assumptions consistently throughout. These are the demands placed on models by AutomationBench, Zapier’s benchmark for business processes. On the leaderboard dated 9 September 2026, Astra successfully completed 41.4% of tasks, compared with 28.77% for GPT-5.6 Sol, with both models using their highest reasoning setting. A task counted as successful only when all required conditions were met. That amounts to roughly 13 more successfully completed tasks out of every 100. For businesses, this is an opportunity to test whether AI can complete more stages without employee assistance. It is worth reviewing tasks where an employee repeatedly prompts the model to take the next step, supplies more data or transfers the output to another application. Each intervention takes time and reduces the benefit of automation. Astra gives businesses reason to test which of these stages it can now handle together. The 41.4% result also shows how demanding these tasks remain. The level of autonomy should be determined by the model’s performance within the company’s own process. 2.2 Automating more tasks in existing business software Companies use many applications introduced at different stages of their development. AI’s access to these resources plays a significant role in how useful it can be. When a model can navigate an application’s interface effectively, businesses can consider using it for tasks that employees currently perform through forms, buttons and menus. In ScreenSpot-Pro, a test of locating elements in screenshots, Astra achieved 92.7% accuracy, compared with 76.9% for GPT-5.6 Sol. The test covers complex applications with densely packed controls. The result measures one specific skill needed to operate software. This improvement gives businesses reason to consider automating more tasks in the applications they already use. Automation planning should include tasks that currently require employees to navigate software manually. Better interface recognition could help AI carry them out when equipped with the appropriate tools and permissions. Cost-effectiveness will depend on the success of the entire operation, including reading data, making changes and checking the result. Accurately locating a button is one requirement for completing that operation successfully. 2.3 Detecting errors that require connecting information from several places Business tasks are often difficult because of the connections between pieces of information. A change to one agreed detail can affect subsequent decisions, documents and team activities. The person responsible for the overall task needs to identify these dependencies and assess their consequences. In CodeRabbit’s evaluation, Astra’s advantage was particularly clear when detecting bugs that required examining several code files. The model detected 57.1% of labelled bugs, compared with 47.6% for GPT-5.6 Sol. Across the full evaluation, the difference was smaller: 61.3% versus 59.0%. The greatest improvement was therefore seen in the more challenging cases. For those overseeing AI adoption, this offers a useful lesson: evaluate the new generation on tasks where the earlier model missed connections between pieces of information or required extensive corrections. With simple prompts, the difference between models may be small. Materials containing exceptions, interdependent conditions and information spread across several places can reveal more about Astra’s usefulness. CodeRabbit documents this improvement in software analysis. Establishing whether similar gains apply to documentation, reports or business agreements requires testing on the company’s own materials. This is a useful way to assess the model in organisations where employees spend considerable time working out how different pieces of information affect one another. 2.4 When does better AI performance translate into savings? Even a small improvement in quality can matter when a task is repeated hundreds of times a month. If employees spend less time correcting outputs and helping AI finish tasks, the team gains time back. The scale of that benefit depends on how often errors occur, how long they take to correct and the consequences of those mistakes. Astra’s results justify reassessing applications of AI that previously proved too unreliable or required too much supervision. A company can revisit a shelved idea and test it on a set of real cases, including more difficult ones. Three measures matter in this assessment: the proportion of tasks completed correctly, employee time spent on checks and corrections, and the total cost of handling each case. These indicators help establish whether the model’s better performance benefits the team as a whole. Astra’s progress could therefore make tasks that were previously too costly to automate economically viable. Tasks that have required constant human assistance are worth testing again, particularly when they are frequent, time-consuming and have a clearly defined outcome. How can your business put GPT-6 Astra to use? Our AI consulting services help you identify the tasks where improvements would deliver the greatest business benefit and plan the implementation. At TTMS, we work with clients to analyse the process, identify the data and integrations required, and agree on how to measure results. We combine consulting with designing AI solutions and integrating them with business systems. We were the first company in Poland to obtain accredited ISO/IEC 42001 certification for our artificial intelligence management system. For clients, this confirms that our practices for risk assessment, project documentation and AI oversight have been reviewed by independent auditors. Tell us about a task you would like to improve. Together, we will explore how AI could help and where to begin. Talk to TTMS about AI consulting for your business. GPT-6 Astra in business: frequently asked questions What does GPT-6 Astra change for businesses already using AI? GPT-6 Astra gives businesses reason to test whether AI can handle a larger share of a task with less employee assistance. Comparisons with GPT-5.6 Sol show improvements in completing tasks across several applications, locating on-screen elements and analysing dependencies in code. It is therefore worth revisiting processes where the team frequently corrects outputs or guides the model step by step. The benefit comes when better performance reduces the work needed to achieve a correct result. Can GPT-6 Astra work in business systems without building a new integration? In some cases, yes, provided its environment gives Astra tools to operate a browser or computer. The model can then interact with applications through forms, menus and buttons. It needs access to the system and the appropriate permissions. Whether this approach is suitable depends on the application and the task. For frequent, repetitive operations, it is worth comparing it with an API integration, which allows software systems to exchange data directly. Which tasks are worth trying with AI again if earlier automation attempts failed? Start with tasks where the earlier model lost track of the steps, confused interface elements or missed connections between pieces of information. These are areas where the tests discussed in this article show Astra’s progress. Use the same materials and assessment criteria for the new trial, including cases that previously proved difficult. If outdated data, conflicting instructions or a lack of application access caused the original failure, those issues also need to be resolved. How can you decide what GPT-6 Astra can do independently and what needs employee approval? The level of autonomy should depend on the consequences of a mistake, whether an action can be reversed and the results of tests using company data. Drafting a document or organising a copy of a dataset allows the output to be checked before use. Sending a proposal, changing commercial terms or deleting records may require prior approval. Process instructions should clearly define when AI can proceed, when it should request a decision and when it should stop because information is missing. How can you tell whether GPT-6 Astra saves your business money after checks and corrections are included? Compare the total cost of completing the same set of tasks with and without Astra. Include data preparation, tool use, reviewing outputs, making corrections and any cases that employees have to redo. A useful measure is the cost per successfully completed task that meets the required quality standard. Base the assessment on a representative set of tasks, including exceptions. This will help establish whether the savings hold up in day-to-day work.
Read