Home •Blog

TTMS Blog

TTMS experts about the IT world, the latest technologies and the solutions we implement.

Sort by topics

Clear all filters

Search results for the term: “TEAMS”

RAG for Chatbots using CrewAI: Notes from a TTMS Tech Talk

RAG for Chatbots using CrewAI: Notes from a TTMS Tech Talk

Tech Talk is an internal series of technology sessions for TTMS employees, where we share knowledge and project experience. We demonstrate tried-and-tested tools, discuss challenges we have encountered and explain the solutions that have helped us in our work. Topics include artificial intelligence, data analytics, Salesforce, AEM and project management. Presentations are followed by time for questions, discussion and sharing ideas. During the session “RAG for Chatbots Using CrewAI”, held on 17 September, Jakub Kraśniewski, Senior AI Developer at TTMS, discussed improvements to a chatbot using a client’s documentation. He presented the challenges involved in preparing and retrieving information, the solution implemented and the approach to evaluating answer quality. The project involved a company in the education sector whose customers were preparing for a certification exam. The chatbot was intended to help them find information about registration, exam procedures, grading and appeals, thereby reducing the support team’s workload. It used several hundred pages of publicly available PDF documents, mainly in English. The team needed a way to retrieve relevant information from these materials while meeting a response time requirement of around 4 to 5 seconds. 1. How does RAG help a chatbot use company knowledge? Jakub began the presentation by explaining how RAG (Retrieval-Augmented Generation), a method of generating answers using retrieved source material, works. The system finds information relevant to the user’s question and passes it to a language model as context for the answer. In this project, the retrieved material consisted of passages from documentation describing exam rules and procedures. After extracting text from the documents, the system divides it into smaller chunks. An embedding model (an AI model that represents semantic features of text as numbers) converts these chunks into vectors stored in a database. The user’s question is processed in the same way. Comparing these representations allows the system to retrieve passages that are semantically related to the question. Jakub emphasised that the team is responsible for the quality of the material passed to the model. This involves checking whether the text was extracted correctly, whether the chunking preserved the necessary context and whether retrieval provides information useful for answering the question. The challenges the team encountered in the CrewAI-based solution demonstrated the importance of these steps. 2. What made information retrieval difficult in the CrewAI project? After explaining the basics of RAG, Jakub shared his experience from a project using CrewAI, a framework for building AI agent-based systems. He discussed three problems encountered in the configuration used: overly large text chunks, the absence of an additional relevance assessment and errors in PDF text extraction. 2.1 Overly large document chunks In the configuration Jakub described, text was split into chunks of 4,000 characters. The system retrieved five such chunks for each question, passing up to approximately 20,000 characters of source material to the model. A large chunk can contain information on several different topics, making it harder to match it to a specific question. The model generating the answer must then select the relevant information from the supplied content. In this project, the chunking approach therefore needed to be adapted to the structure of the documents and users’ questions. 2.2 No additional assessment of search result relevance Jakub pointed out that the configuration lacked reranking, which involves reassessing and reordering search results according to their usefulness for answering the user’s question. The system can first retrieve a larger number of passages, then assess them further to select those most useful for preparing an answer. Jakub presented this method as a potential improvement whose value should be evaluated by checking both answer quality and response time. 2.3 Incorrect text reading order in multi-column PDFs Another problem involved document text extraction. The tool read multi-column PDFs row by row, merging content from adjacent columns. This disrupted the order of sentences and made subsequent information retrieval more difficult. The resulting text was then split into chunks. The error therefore originated during data preparation and affected the subsequent stages of document processing. This example showed why evaluating RAG quality should begin with comparing the extracted text against the source document. 3. How does response time affect the choice between Classic RAG, Agentic RAG and Graph RAG? A key project requirement was a response time of around 4 to 5 seconds. Jakub discussed three RAG approaches in terms of data preparation costs, the ability to evaluate their operation and the time needed to handle a question. Approach How it works, as discussed during the session What to consider when choosing Classic RAG Retrieves passages from a knowledge base, optionally reranks them and passes the context to the model. Document chunking quality, retrieval relevance and the amount of context provided. Agentic RAG An agent selects tools and a retrieval method, running additional queries as needed. The ability to adapt retrieval to the question, along with the time and cost of additional operations. Graph RAG Retrieval uses a knowledge graph describing entities found in the source material and the relationships between them. The effort required to build and maintain the graph, and how useful the relationships are for answering users’ questions. In an agentic approach, the model can use several tools, such as vector search, keyword search or filtering by metadata describing the document. Additional steps allow the system to expand its search for information, while their number and sequence affect response time. In the graph-based approach presented, some of the work takes place when building the knowledge base. Entities and the relationships between them are extracted from the text. This mechanism also underpins GraphRAG as described by Microsoft. Jakub highlighted the costs of this preparation and the difficulty of manually analysing a complex graph. In this project, the response time requirement favoured further development of classic RAG. The team focused on document chunking and context selection. 4. How does hierarchical document chunking help preserve context? The solution organised the material into three connected levels: pages, paragraphs and sentences. The system retained information about which paragraph each sentence belonged to and which page contained that paragraph. Content was represented in the vector database at different levels of detail. This allowed retrieval to identify both individual sentences and larger passages containing the required information. According to Jakub, the additional cost of storing and processing these representations was acceptable given the volume of material in the project. Finding a relevant sentence made it possible to retrieve its entire paragraph and provide the model with broader context. The system could also retrieve the whole page when needed. Suppose a user asks about the deadline for appealing an exam result. The system finds a sentence specifying the deadline, then retrieves the entire paragraph explaining when the appeal period begins and how to submit an appeal. This allows the model to account for these conditions in its answer. 5. How can you evaluate RAG quality using your own data? In the final part of the presentation, Jakub emphasised the importance of a benchmark, a set of tests used to compare different versions of a solution. He discussed checking retrieval results against information labelled by a human and using a language model to evaluate answers. In practice, it is useful to assess two stages separately. The first concerns retrieval: did the system return a passage containing the required information? The second concerns the answer: did the model use the supplied material correctly? This distinction helps identify which stage needs improvement. In additional information shared after the session, Jakub clarified the testing method and results. The test set included questions covering the full scope of the documentation, along with real user questions collected anonymously during a prototype launch at the beginning of the year. Answer accuracy increased from around 70% to around 98%, an improvement of approximately 28 percentage points. This result applies to the internal test conducted in this project. According to Jakub, the solution also maintained a fast response time. When he shared these details, the chatbot had completed internal testing, and the company planned to make it available to a subset of customers. Reducing the support team’s workload and making information easier to access remained deployment goals. Assessing whether those goals have been achieved requires data from actual use. The embedding model is another component to evaluate. Jakub noted that its selection should take into account the language of the source material and its performance on the team’s own dataset. The choice of this model affects which passages the system retrieves before it begins generating an answer. 6. What can companies implementing a chatbot learn from this experience? The project shows how specific requirements guide RAG development. The expected response time helped narrow down the choice of solution, document analysis revealed problems with text extraction and chunking, and an internal test made it possible to assess the impact of the changes. When planning a similar implementation, it is worth addressing five areas: Source material: check whether document text extraction preserves meaning and reading order. Document chunking: adapt chunk size and the connections between chunks to the structure of the material. User questions: prepare a test set that reflects the tasks the chatbot is intended to support. Response time: establish expectations and account for them when comparing approaches. Quality assessment: check both the relevance of the retrieved information and how it is used in the answer. Let’s talk about AI in your company TTMS is home to experts who, like Jakub, combine technical knowledge with experience from client projects. During Tech Talks, they share solutions tested in practice and apply what they have learned to subsequent implementations. Are you planning a chatbot that uses your company’s documentation, or looking to improve the answer quality of an existing tool? Let’s talk. We will review your materials, users’ needs and business goals to recommend an appropriate way to use AI. Contact the TTMS team! What documents can a RAG chatbot use as a knowledge base? A RAG chatbot can use company policies, product manuals, procedures, FAQs and other materials containing information relevant to its users. Sources may include PDFs, Word documents, website content and knowledge base articles, depending on the integrations available. Scanned documents require optical character recognition (OCR) to turn images of text into searchable content. Tables, diagrams and complex layouts may need additional processing to preserve their meaning. Before adding documents, check that they are accurate, current and approved for the intended audience. Clearly structured materials help the system retrieve information and provide useful context for its answers. How do you keep a RAG chatbot’s knowledge base up to date? Keeping a RAG chatbot up to date requires a process for detecting and processing changes in its source materials. Depending on business needs, updates can run on a schedule or be triggered when a document is added, edited or removed. The system then updates the searchable content and its associated representations, such as embeddings. Version information and effective dates help distinguish current guidance from older material. Deleted or superseded documents should also be removed from active search results, and cached answers may need refreshing. Assigning an owner to each content area helps ensure that someone remains responsible for the information the chatbot uses. Can a RAG chatbot provide sources for its answers? Yes, a RAG chatbot can include links, document titles, page numbers or quoted passages alongside its answers. This requires the system to preserve source information when processing documents and connect retrieved passages to the response. Useful citations let users open the relevant material and check the context for themselves. The system should also verify that each citation supports the claim it accompanies. A source link alone provides no guarantee that an answer accurately reflects the document. During testing, teams should check both answer quality and citation accuracy, including whether users can access the referenced material. How can a RAG chatbot respect access permissions for company documents? A RAG chatbot can use the signed-in user’s identity and access rights to determine which documents it may retrieve. Permission checks should happen before restricted content reaches the language model. The same controls need to cover document previews, citations and any cached responses that contain protected information. When access rights change in a source system, those changes must also be reflected in the chatbot’s retrieval process. Teams should test the solution using accounts with different roles, including users with limited access. These checks help confirm that each person receives answers based on information they are authorised to view. What should a RAG chatbot do when it cannot find an answer? When the available documents provide insufficient information, a RAG chatbot should clearly explain that it cannot answer reliably from its sources. It can ask a clarifying question if the request is ambiguous or suggest a related document that may help. For questions requiring further assistance, it can direct the user to the appropriate team or support channel. The system needs explicit rules for handling incomplete, conflicting or missing information. Testing should include questions whose answers are absent from the knowledge base, so the team can assess this behaviour. Reviewing unanswered questions can also reveal gaps in company documentation and priorities for future updates.

Read
ChatGPT for financial services: what does combining GPT with professional data sources offer?

ChatGPT for financial services: what does combining GPT with professional data sources offer?

On 10 September 2026, OpenAI announced ChatGPT for Financial Services, a solution combining GPT-6 Astra with professional financial data sources and tools for preparing analyses. The product was developed in collaboration with Morgan Stanley and Evercore. It is designed for financial institutions, with an initial focus on investment banking and equity research. It allows analyst teams to find data, perform calculations and prepare client materials in one place. This could reduce the time spent gathering information and transferring it between tools. In this article, you will learn: what data and features ChatGPT for Financial Services offers, what preparing a company analysis with GPT could look like, why metric calculations and data sources need to be checked, which stages require an analyst’s review, how to assess whether implementation is worthwhile for your company. How does ChatGPT for Financial Services support analysts? Materials published by OpenAI and its data providers describe several specific use cases: Comparing companies. Daloopa, a provider of company financial data, makes selected data and metrics available for comparing business performance. Source references help analysts verify where the figures come from. Finding companies that meet specific criteria. Daloopa also describes searching for companies by business activity or geographical region. The resulting list can provide a starting point for further market analysis. Preparing client materials. OpenAI describes creating valuation models, research notes and presentations using company templates for Excel, Word and PowerPoint. ChatGPT for Financial Services provides access to selected data from Daloopa, PitchBook, LSEG News and Crunchbase. OpenAI is also developing integrations that will allow institutions to use data covered by their existing subscriptions. These include S&P Capital IQ, LSEG, MSCI, Dow Jones Factiva and Moody’s. From data to company analysis: five steps in the workflow Let’s walk through preparing a comparison of two industrial companies ahead of a client meeting. The analyst needs to assess profitability, explain material differences and prepare a short note with a results table. Using fictional data, we will show what to check during a pilot, from selecting information to approving the final material. 1. Defining the question and scope of the comparison First, we establish which period to compare: the last full year, a six-month period or the trailing twelve months. We also check whether the figures cover the entire corporate group or an individual company, which currency they use and how each metric was calculated. In our example, we use consolidated data for both groups for the same calendar year. Amounts are stated in millions of euros. Before comparing results, we need to check the start and end dates of the reporting periods. One company’s financial year may end in December, while another’s ends in March. The documentation for the US SEC’s EDGAR database also highlights these differences. The analyst must then decide how to account for the mismatch and whether additional data is needed. The agreed approach should be recorded in the instructions for the AI and included with the completed analysis. This gives the model clear guidance and helps the reviewer understand which data was compared and why. 2. Gathering data and identifying its sources For each important figure, record the company and period it relates to, the units used and how it was calculated. A reference to the specific table or explanatory note in the report is also needed. Keeping the source document and its retrieval date makes it easier to review or update the analysis later. According to OpenAI’s description, ChatGPT for Financial Services lets users locate specific tables and document passages, highlighting the information used in the analysis. In our example, we compare EBITDA, or earnings before interest, taxes, depreciation and amortisation. The reviewer should be able to trace a reported value back to the company’s report and check how it was calculated. This also allows them to confirm that the figure covers the correct period and scope of operations. 3. Aligning definitions before comparing margins Companies may report adjusted EBITDA that excludes selected costs. These adjustments increase the value of the metric. Before comparing profitability, it is therefore necessary to check which adjustments have been applied. The US SEC also highlights differences in how individual companies calculate financial measures. Let’s look at two fictional companies. We assume that both calculate EBITDA before adjustments using the same principles. Company A then adds back EUR 4.59 million in costs that it excludes when calculating adjusted EBITDA. As a result, the metric rises from EUR 27.57 million to EUR 32.16 million. Company B has no such costs, so its figure remains unchanged. Illustrative example. Consolidated data for the same calendar year. Amounts are stated in EUR million. Item Company A Company B Revenue 229.73 183.78 EBITDA before adjustment 27.57 23.89 Costs excluded when calculating adjusted EBITDA 4.59 0.00 Adjusted EBITDA 32.16 23.89 EBITDA margin before adjustment 12% 13% Adjusted EBITDA margin 14% 13% After the adjustment, Company A’s margin is 14%, exceeding Company B’s margin of 13%. Before the adjustment, Company B has the higher margin: 13% compared with 12%. In this example, the treatment of costs determines which company has the higher EBITDA margin. The analyst should therefore check which costs make up the EUR 4.59 million adjustment and whether they also occurred in previous years. This helps them assess whether excluding these costs is justified for the analysis being prepared. They can also present both scenarios and explain the difference to the client. AI can help gather data and recalculate margins, while the expert assesses whether the adjustment is justified and how it affects the conclusions. 4. Verifying calculations in the spreadsheet In our example, simply divide EBITDA by revenue: 27.57 ÷ 229.73 gives a margin of approximately 12%, while 32.16 ÷ 229.73 gives approximately 14%. The displayed amounts are rounded; the spreadsheet should retain full precision for its calculations. More complex analyses may require currency conversion, alignment of reporting periods or the preparation of several forecast scenarios. The reviewer should be able to trace each of these steps. It is therefore worth asking AI to create a spreadsheet in which source data, assumptions and formulas are clearly separated. The analyst can then check the calculations and see how changing a single value affects the result. To test the spreadsheet, you can halve the adjustment, reducing it from approximately EUR 4.59 million to EUR 2.30 million. Company A’s adjusted EBITDA should then be approximately EUR 29.86 million, with a corresponding margin of 13%. These amounts are rounded for presentation; the spreadsheet should calculate the change using unrounded values. After making this change, check the comparison table and the commentary on the results as well. Both companies would now have the same margin, so the conclusion that Company A has a higher margin would need updating. This is a simple way to assess whether the calculations and accompanying text remain consistent. 5. Preparing client materials and reviewing conclusions The completed analysis can be presented in the company’s preferred format. According to OpenAI’s description, an administrator can share Excel, Word and PowerPoint templates with the team for the tool to use when creating documents and presentations. In our example, the client should receive a results table and a short explanation of how the cost adjustment affects the margin comparison. The expert reviewing the material checks whether the conclusions match the calculations and answer the client’s question. If anything needs clarification, they can request a further explanation or another version of the analysis. Time measurements should also include reviewing the material and making corrections before approval. How can you protect data and preserve a record of the analysis? When preparing a client analysis, the team may use public reports, paid databases and confidential documents. It is necessary to establish who can access this information, where it will be stored and who can receive the finished material. According to the ChatGPT Work security documentation, business data is encrypted and is not used to train models by default. Data retention periods, processing locations and the scope of recorded activity depend on the settings and connected services. Before implementation, check which user and tool actions are logged and which records can be exported. During the pilot, keep the source documents, successive versions of the spreadsheet and the final material, together with a record of who approved it and when. Then check whether this documentation allows you to reproduce the calculations and trace the approval of the analysis. Separately, verify whether system logs allow user and tool activity to be traced to the extent required by the company. How can you assess whether implementation is worthwhile? Start by choosing a task the team performs regularly, such as updating a company comparison after quarterly results are published. Before testing, measure how long this analysis takes using the existing method and define its quality requirements. These findings will provide a baseline for comparison with AI-assisted work. The time needed to review and correct the analysis must be added to the preparation time. OpenAI highlights this in its guidance on assessing the business value of AI, also recommending that implementation and ongoing usage costs be included. In practice, it is worth comparing: Metric What does it tell us? Time from starting the task to approving the analysis Does the client receive the finished material sooner? Include waiting time between stages. Total time spent by the analyst and reviewer Does the team spend fewer hours on preparation, review and corrections? Number of errors affecting the results or conclusions Does the analysis meet the same quality requirements as the existing approach? Accuracy and completeness of source references Can the origins of key figures and information be verified? Time needed to update the analysis How efficiently can new data be incorporated and the calculations and conclusions that depend on it be updated? Cost per approved analysis What is the cost of the finished material, including team time, the tool, data and the share of implementation and maintenance costs allocated to that analysis? The test should cover several tasks of varying difficulty. Define the assessment criteria before it begins. Someone performing the same analysis for a second time already knows the data and some of the answers, which may shorten the time needed. It is therefore worth using comparable tasks and varying the order in which participants work with AI and with the existing method. If AI saves time, check how the team used it. They may have prepared more analyses, responded to clients sooner or reduced overtime. The implementation assessment should show separately how the time saved was used and whether company spending decreased, and by how much. When is it worth starting a pilot? Consider a pilot if the team regularly gathers data from multiple sources and updates similar analyses. Choose a task that takes analysts a significant amount of time, such as comparing data from company reports. Assign a person to lead the pilot and experts to review the results. If the team only occasionally analyses a few annual reports, check whether tools already approved for use within the company are sufficient. Where data retrieval and calculations are already automated, identify a specific task that the new tool could improve. OpenAI makes the product available to financial institutions that meet its access requirements and directs interested companies to its sales team. Pricing, detailed terms and availability for a particular institution in Poland must be confirmed with the provider. Availability information. The implementation decision should be based on the pilot results: the quality of the analyses, the time needed to prepare and review them, and the total cost of the work. The test will also show whether the tool provides access to the data the team needs. Want to explore where AI could improve analysts’ work in your organisation? Talk to the TTMS team about choosing a task for a pilot, connecting the necessary data sources and assessing the results. How does ChatGPT for Financial Services differ from analysing reports in ChatGPT? ChatGPT for Financial Services provides access to selected professional financial data directly within the tool. It also supports references to specific tables and document passages, as well as the preparation of materials using company templates. When assessing its suitability for a team, check whether the available sources cover the companies, periods and metrics the team needs. Does ChatGPT for Financial Services require separate financial data subscriptions? Selected datasets are included in the product. These cover some of the information supplied by the providers named by OpenAI. The company is also developing integrations intended to let institutions use data covered by their existing subscriptions. Before purchasing, confirm which data is included in the offering and which requires additional access rights. Can ChatGPT for Financial Services be used to analyse companies listed on the Warsaw Stock Exchange? This depends on the availability of data for individual companies. Check whether the tool provides their financial statements, relevant metrics and historical data. The launch announcement alone does not confirm full coverage of the Warsaw Stock Exchange. The best way to assess the product’s suitability is to test it on several companies the team regularly analyses. What should you do if data from ChatGPT differs from the figures in a company’s report? Start by comparing the sources, reporting periods, units and definitions of the metrics. A discrepancy may arise, for example, from using standalone rather than consolidated financial data, or from including EBITDA adjustments. Also check whether the company has published an updated report. The analyst should explain the discrepancy and document which value they used and why.

Read
Automation Best Practices in Software Testing for 2026

Automation Best Practices in Software Testing for 2026

Software release cycles keep shrinking, and testing teams are expected to keep pace without sacrificing reliability. Automation has become a core part of modern software delivery, but writing more scripts does not automatically lead to faster feedback, stronger coverage, or more dependable releases. Teams getting real value from automation treat it as an engineering discipline integrated with the wider QA and software delivery process. This article covers the practices that make automation sustainable in 2026, including test prioritization, methodology and tool selection, maintainable script design, CI/CD integration, continuous maintenance, and code ownership. 1. Why Test Automation Best Practices Matter More in 2026 Applications increasingly depend on connected services, frequently changing interfaces, and shorter delivery cycles. Automated testing may cover web applications, mobile products, APIs, integrations, and business-critical user journeys, making the way automation is planned, implemented, and maintained as important as the number of tests in the suite. Without clear standards, automation can become another source of delivery friction. Unstable tests slow down CI/CD pipelines, outdated scripts lose alignment with changing requirements, and duplicated coverage increases execution and maintenance effort without producing better evidence. As discussed in our guide to AI end-to-end testing, sustainable automation requires clear ownership, traceability, regular maintenance, and a deliberate connection between requirements, test intent, execution, and results. 2. A Decision Framework for What to Automate in Software Testing Not every test belongs in an automated suite. Teams need a repeatable way to identify scenarios where automation will provide lasting value rather than create additional maintenance work. 2.1 Scoring Test Cases by Frequency, Stability, and Cost A practical scoring model evaluates each candidate across several dimensions: how often the test runs, how stable the underlying feature is, how important the workflow is to the business, and how much effort the test will require to automate and maintain. Strong candidates are usually repeatable scenarios with clear expected results, stable preconditions, reliable test data, and a meaningful impact on release confidence. A frequently executed test may be valuable, but frequency alone is not enough. Teams should also consider whether the scenario can be executed consistently and whether its expected outcome can be verified without subjective interpretation. 2.2 Keep Judgment-Based Testing Manual Exploratory testing, first-impression usability reviews, visual assessments, and scenarios that depend heavily on human interpretation are usually better handled manually. Automating these activities can remove the flexibility and observation that make them valuable. Rarity alone should not exclude a test from automation. An unusual scenario may still deserve automated coverage when a failure could interrupt a critical process, compromise data, or create significant operational risk. The decision should reflect business impact as well as execution frequency. 2.3 Prioritize Business-Critical User Paths After unsuitable candidates have been removed, rank the remaining tests by business impact and regression risk. Authentication, checkout, account access, approvals, and core transaction flows are common priorities because failures can prevent users from completing essential tasks. Starting with these workflows allows the automation suite to protect the parts of the application that matter most. Lower-impact scenarios can be added later when their expected value justifies the development and maintenance effort. 3. Match the Automation Approach to the Application and Team Once teams have identified the right candidates for automation, they need to choose an approach that fits the application, delivery model, and available skills. The decision should balance execution speed, maintainability, technical control, accessibility for QA, and the effort required to integrate automation with the existing development workflow. 3.1 Use the Test Pyramid to Balance Feedback and Coverage The test automation pyramid remains a useful starting point. Fast unit tests usually form the broadest layer, integration and service-level tests verify interactions between components, and a smaller set of end-to-end tests validates complete user journeys. The exact proportions should reflect the application architecture and its risks. A system built around multiple services may need stronger integration coverage, while a business application may require more end-to-end validation of critical workflows. The goal is to detect problems at the lowest practical level while retaining enough end-to-end coverage to confirm that essential user journeys work as expected. 3.2 Choose Between Code-Based, Codeless, and Hybrid Automation Code-based automation provides direct control over test architecture, integrations, reusable components, and repository conventions. It is often appropriate when a team has strong engineering skills, complex testing requirements, or an established automation framework. Codeless and low-code approaches reduce the amount of scripting required to define common test scenarios. They can make automation more accessible to manual testers and domain specialists, although teams should still evaluate how the platform handles complex logic, maintenance, version control, and code ownership. A hybrid approach combines accessible test definition with standard code-based automation. QA professionals can define and review test intent, while automation engineers maintain technical standards and review the resulting code. This model can reduce handovers without limiting the team to a proprietary execution format. 3.3 Evaluate Technical Fit and Long-Term Ownership The right methodology depends on what the team is testing and who will maintain the automation. Complex backend behavior and service integrations may require direct technical control, while repeatable web user journeys may be suitable for higher-level automation. Teams should also consider where generated automation will run, whether the output can be reviewed in the existing repository, how results return to the QA workflow, and whether the chosen approach supports the organization’s deployment and security requirements. Tool popularity matters less than alignment with the team’s application, skills, governance model, and long-term maintenance responsibilities. 4. Design Test Automation for Maintainability Script and framework design have a major influence on long-term reliability, regardless of the selected tool. Maintainable automation separates reusable technical components from test intent, follows consistent repository conventions, and makes failures easier to understand and repair. 4.1 Using Stable Locators and Resilient Element Selection Locators based on visual position, generated identifiers, or deeply nested DOM paths can break when the interface changes. Where possible, teams should use selectors based on stable and meaningful attributes, such as dedicated test identifiers, accessible roles, labels, or other elements that reflect how users interact with the application. A locator should be both stable and specific enough to identify the intended element. The goal is not to eliminate maintenance entirely, but to reduce unnecessary failures caused by implementation details that are unrelated to the behavior being tested. 4.2 Apply Reusable Design Patterns Patterns such as the Page Object Model can separate interactions with the application from the business logic being verified. When an interface element changes, the team can update the relevant reusable component instead of modifying every test that uses it. The appropriate structure may also include component objects, shared fixtures, helper functions, reusable actions, and project-specific abstractions. Teams should select patterns that match the application architecture and apply them consistently across the repository. 4.3 Define Clear Assertions Every automated scenario should include an explicit and observable expected result. A test that performs a sequence of actions without verifying the outcome may pass even when the underlying business process is not working correctly. Assertions should confirm the intended behavior rather than incidental implementation details. Clear expected results also make test cases easier to review, automate, diagnose, and trace back to the original requirement. 4.4 Manage Test Data as a Dedicated Discipline Test data should have clear ownership and a repeatable setup process. Data creation, seeding, reuse, protection, and cleanup should be considered when the test is designed rather than added after the script has already been implemented. Teams should also avoid hidden dependencies on data left behind by earlier executions. Predictable test data makes failures easier to reproduce and reduces false results caused by stale, incomplete, or conflicting records. 4.5 Keep Tests Independent and Isolated Automated tests should not depend on the order in which other tests are executed. Each test should begin from a known state and establish the preconditions required for its own scenario. Isolation can be implemented through fixtures, controlled data setup, separate browser contexts, API-based preparation, environment resets, or cleanup procedures. When one test fails, that failure should not create misleading results elsewhere in the suite. 4.6 Follow Repository Conventions Automation code should follow the same structural and quality standards as the rest of the project. Consistent naming, fixtures, helpers, hooks, error handling, formatting, and review rules make generated and manually written tests easier to understand and maintain. Repository alignment becomes particularly important when automation code is generated rather than hand-written. Generated automation should fit the existing framework rather than introduce a parallel structure that the engineering team must maintain separately. 5. Integrate the Right Tests at the Right CI/CD Stage Automated tests provide the greatest operational value when they are integrated with the delivery pipeline and return timely, actionable results to the people responsible for the change. The goal is not to run every test after every update, but to apply the right level of validation at each stage. 5.1 Match Test Scope to the Pipeline Stage A staged pipeline balances feedback speed with test depth. Fast unit tests and static checks can validate individual changes early, while integration tests and a relevant regression subset can provide broader evidence during pull or merge request review. More extensive regression testing can run after changes are merged, before deployment, on a schedule, or when the risk of a release justifies wider coverage. Smoke tests serve a different purpose. They verify that a deployed application is available and that its most critical user paths remain operational. They should complement, rather than replace, deeper regression testing. The exact structure and runtime expectations should reflect the application architecture, infrastructure, release cadence, and business risk. Teams should define their own quality gates based on the type of change and the evidence required before it can move forward. 5.2 Select Tests Based on Change and Risk Running the full suite for every code change can create unnecessary delays as automation grows. A more sustainable approach selects tests based on the components affected by the change, related requirements, business-critical workflows, previous execution results, and known areas of risk. Test selection should remain transparent. Teams need to understand why a test was included or excluded and should be able to expand the scope when a change has wider implications than the initial analysis suggests. This staged, risk-based structure is also the foundation that AI-driven test selection and prioritization build on. If you’re looking at how AI fits into this part of the pipeline specifically, see our guide, AI in Software Test Automation: 2026 Guide. 5.3 Return Actionable Results to the Workflow Test execution should produce more than a pass or fail status. Results should identify the affected scenario, provide enough evidence to investigate a failure, and remain linked to the relevant requirement, test case, code change, and execution record. Returning results to the pull or merge request helps engineering teams review automation alongside the application change. Returning them to the QA layer and requirement source gives QA teams traceability from test intent through code to execution. This closes the feedback loop and supports informed release decisions. 6. Treat Test Maintenance as a Continuous Engineering Process Flaky tests can quickly undermine confidence in an automation suite. When the same test passes and fails without a relevant application change, teams may begin treating failures as noise. Rerunning the test can unblock a pipeline temporarily, but it does not resolve the underlying problem. 6.1 Diagnose the Root Cause of Flaky Tests Flakiness can originate in the test code, application, test data, execution environment, or an external dependency. Common causes include fixed delays, missing synchronization, unstable locators, shared state between tests, conflicting test data, asynchronous application behavior, network variability, and resource constraints in the execution environment. The corrective action should match the source of the problem. Timing failures may require condition-based waits rather than longer fixed delays. Shared-state failures call for stronger isolation and controlled data setup. Selector failures may require more stable locators or reusable application abstractions. Failures caused by external dependencies may require controlled test environments, mocks, or other ways to reduce unnecessary variability. Teams should capture enough information to distinguish a product defect from a test defect or an environment failure. Execution traces, logs, screenshots, videos, network activity, environment details, and the affected test step can make intermittent failures easier to reproduce and diagnose. 6.2 Review and Prune the Test Suite Regularly Automated tests should be reviewed throughout their lifecycle. Product changes can make tests obsolete, duplicate existing coverage, or reduce their business value. Keeping every test indefinitely increases execution time and maintenance effort without necessarily improving release confidence. A defined maintenance cadence should include reviewing flaky tests, retiring obsolete scenarios, consolidating duplicate coverage, and reassessing tests whose business impact or technical stability has changed. The frequency should reflect the size of the suite, release cadence, application risk, and volume of product changes. Each test should have a clear outcome after review. It may be repaired, rewritten, moved to a more appropriate test layer, temporarily isolated with an assigned owner, or removed when it no longer provides useful evidence. 6.3 Track Execution Health and Maintenance Effort Pass rate alone does not show whether an automation suite is healthy. Teams should also monitor recurring failures, rerun frequency, execution duration, maintenance effort, obsolete tests, duplicated coverage, and the time required to identify and repair a failure. These signals help distinguish a growing automation suite from a sustainable one. Adding tests increases coverage only when the team can understand the results, maintain the assets, and trust the evidence produced by each execution. Maintenance decisions should also remain traceable. Teams need visibility into what changed, why a test was updated, how the change was reviewed, and whether the revised test still validates the original requirement. We will discuss this lifecycle-based approach in our article: Best QA Practices in Software Testing. 7. Treat Automation Code as Production Code Automation code should follow the same engineering standards as application code. Tests need clear ownership, version control, consistent repository conventions, review rules, and quality gates. Without these practices, scripts become difficult to understand, update, and trust as the application evolves. Ownership should cover both test intent and technical implementation. QA professionals can define the business scenario, expected result, and required coverage, while automation engineers or developers verify that the resulting code follows project conventions and can be maintained within the existing framework. Whether automation is written manually or generated with the help of AI tooling, it should enter the repository through a standard pull or merge request. Reviewers should be able to inspect what the test validates, how it interacts with the application, which reusable components it uses, and whether it introduces unstable dependencies or duplicated coverage. Generated tests should not remain inside a proprietary execution environment when code ownership and portability matter to the organization. Keeping automation in the client’s repository makes changes visible, reviewable, and subject to the same governance as other project assets. Teams should also define who responds when a test becomes unstable or outdated. Clear ownership prevents failed tests from remaining unresolved because QA, development, and automation teams each assume that another group is responsible. 8. Where AI Fits Into This Process AI is increasingly used alongside these practices — drafting test cases from requirements, verifying that a proposed path actually works before code is generated, and flagging tests that are likely to become unstable. Getting this right is less about the automation practices covered above and more about strategy, tooling, and governance, so we cover it separately in depth in AI in Software Test Automation: 2026 Guide. 9. Common Test Automation Mistakes to Avoid Even a technically sound automation program can lose value when teams make poor decisions about scope, ownership, and maintenance. Common mistakes include: Automating every available scenario instead of prioritizing repeatable, stable, and business-critical tests. Treating reruns as a permanent solution to flaky tests rather than investigating their root causes. Allowing obsolete and duplicated tests to remain in the suite without regular review. Selecting a platform based on feature count without evaluating integration, maintainability, deployment, code ownership, and team fit. Generating automation without first verifying that the test scenario works against the actual application. Accepting generated tests without checking whether they reflect the original requirement and intended business behavior. Keeping automation in a proprietary environment when the organization requires reviewable, portable code in its own repository. Managing test scripts outside standard version control, code review, and CI/CD quality gates. Leaving responsibility for test intent, code quality, and ongoing maintenance undefined. Avoiding these mistakes requires more than adding tools or scripts. Teams need a controlled workflow that connects requirements, test design, execution evidence, automation code, review, and maintenance. 10. How Qatana Applies These Test Automation Best Practices Qatana brings these practices together in an agentic test automation platform. It turns Jira or GitLab requirements into automation-ready test cases, verifies the proposed user path through Playwright execution, and generates standard Playwright code delivered through a pull-ready merge request. QA retains control of test intent, while engineering reviews the generated automation through the existing merge request process. Execution results remain linked to the requirement, test case, and code, providing traceability across the workflow. Qatana runs entirely on-premise and supports the organization’s selected LLM, keeping project context, generated code, and test evidence within the customer’s environment. Book a demo to see how Qatana turns requirements into execution-validated Playwright automation. 11. Frequently Asked Questions How much of my test suite should be automated? There is no universal target. Prioritize stable, repeatable, and business-critical scenarios, while keeping exploratory and judgment-based testing manual. What is the difference between a test automation strategy and a framework? A strategy defines what to automate, why, and how success will be measured. A framework is the technical structure used to build, execute, and maintain automated tests. How do I know if my automation ROI is positive? Compare the time and effort saved through automation with implementation, execution, and maintenance costs. If maintenance consistently outweighs the benefits, review the selected tests and automation approach. What’s the most common reason automation programs stall after an initial pilot? Usually a lack of ongoing ownership rather than a tooling problem. Teams automate an initial batch of tests successfully, but without a defined maintenance cadence, clear code ownership, and a repository review process, the suite accumulates flaky and obsolete tests faster than anyone repairs them, and confidence in the results erodes.

Read
How to Prepare Data for Power BI Copilot and Build a Semantic Model for AI

How to Prepare Data for Power BI Copilot and Build a Semantic Model for AI

Copilot can speed up data analysis, visual creation and work with semantic models. It does not replace well-managed data sources, correct relationships or agreed metric definitions. If a model contains several similar sales measures, technical column names and ambiguous relationships, the AI assistant inherits the same problems that already affect report users. The difference is that a natural-language answer may sound convincing even when it relies on the wrong metric. To prepare data for Power BI Copilot, organisations need to improve model quality, provide business context and define how the data may be used. Enabling Copilot alone will not make an inconsistent model unambiguous. The model needs clear names, validated measures, a controlled data scope, appropriate permissions and a test set based on questions that users actually ask. This guide explains how to prepare Power BI data and semantic models for Copilot, how to use Prep data for AI and how to assess readiness before making the solution available to employees. It focuses on implementation and ongoing quality control. The broader capabilities of the assistant are covered in a separate TTMS article about AI and Copilot in Power BI. 1. Power BI model readiness for AI An AI-ready model allows users to ask questions in business language without knowing the technical names of tables and columns. This does not mean that Copilot knows the organisation or can discover every internal rule on its own. It can use the information provided through the model, its metadata, the AI feature configuration and the report context. The quality of this layer determines whether a question about sales is mapped to the official net revenue measure, order value or another similarly named field. Assess model readiness across five areas: the quality and freshness of source data; the accuracy and simplicity of the semantic model; unambiguous business concepts, names and measures; data security, permissions and ownership; a repeatable process for testing Copilot answers. If one of these areas is weak, refining prompts will have limited value. A user may phrase a question more precisely but still cannot know which of three margin measures is official or why one table uses the order date while another uses the invoice date. 2. Why the semantic model determines answer quality The semantic model sits between source data and the report user. It contains tables, columns, measures, relationships, formats, hierarchies and security rules. For Copilot, it is the main source of information about how data is organised and how it should be interpreted. The assistant may use the model schema, relationships, object properties, data types, formats and selected metadata. It should not be expected to infer definitions that carry a specific meaning within the organisation. If an active customer is defined as a customer who purchased something in the past 90 days, that definition should be implemented in the model and used by the official measure. Leaving several plausible interpretations increases the risk that Copilot will select the wrong one. 2.1 An example of an ambiguous model Suppose a model contains measures called Sales, Total Sales, Net Sales and Sales Adjusted. An analyst familiar with the project may understand the differences. A business user and Copilot see four plausible answers to the question, “What were sales last quarter?” The fix is not simply to add an instruction telling Copilot to choose one measure. First, the organisation should agree the official definition, give it a clear name, describe the calculation, hide technical fields and remove unused objects. An AI instruction can then clarify the context, but it should not compensate for disorder in the model. 3. Technical requirements before work begins Before redesigning the model, confirm that the environment meets Microsoft’s current requirements. Copilot availability depends on tenant settings, permissions, the workspace and the assigned capacity. Requirements can vary between Power BI Desktop, the Power BI service and individual Copilot experiences. At the time of writing, Microsoft’s documentation for Copilot in Power BI identifies requirements that include enabling the relevant tenant setting and using supported paid Fabric capacity or Power BI Premium capacity. Model authors also need the appropriate permissions for the workspace and semantic model. Prep data for AI is currently a preview feature. Before implementing it, account for current limitations, including the requirement to enable Power BI Q&A and the connection types supported in Power BI Desktop. Area What to check Why it matters Capacity Whether the workspace uses supported paid Fabric or Power BI Premium capacity Without the required capacity, some Copilot experiences may be unavailable Tenant settings Whether the administrator has enabled Copilot and Azure OpenAI based features for the relevant groups Access should be granted deliberately and in line with organisational policy Permissions Whether the author can edit the model and publish to the target workspace Model configuration requires permissions appropriate to the environment Connection type Whether the connection is supported by the feature used in Desktop or the service Support differs between tools and may change Power BI Q&A Whether Q&A is enabled for the model This is currently one of the requirements for Prep data for AI Licensing and feature requirements should not be copied into an internal procedure and treated as permanent. Microsoft continues to develop Copilot and change the availability of individual experiences. Check the latest documentation and the organisation’s tenant settings before implementation. 4. How to prepare data for Power BI Copilot Data needs to be reliable before it reaches the model. Copilot will not repair missing records, incorrect customer mappings or inconsistent currency codes. It can, however, use faulty data to generate an answer that hides the problem behind a plausible explanation. 4.1 Source data quality Start with data profiling and quality controls. Check completeness, uniqueness, consistency, freshness and compliance with business rules. These controls should be repeatable, not limited to the period before the first publication. In practice, this includes: identifying missing keys and orphaned records; checking for duplicates in dimension tables; aligning time zones, calendars, currencies and units; confirming that status values have the same meaning across systems; assigning data ownership and a process for resolving quality issues; monitoring data freshness and failed refreshes. Every official metric should have an identified source, owner, refresh frequency and calculation rule. This gives the organisation a reference value against which to test a Copilot answer, rather than judging the answer only by whether it sounds reasonable. 4.2 Names that match business language and stable definitions Technical names such as fct_sales_hdr, cust_id and rev_net_adj make the model harder for people to use and make user intent more difficult to interpret. The semantic layer should use names that match the language of the organisation, such as Sales, Customer, Net Revenue and Sales Region. Readable names alone are not enough. Margin remains ambiguous if the organisation uses margin amount, margin percentage, planned margin and adjusted margin. Each measure needs a precise name and definition. If one measure is official, the model should reflect that status through naming, its description, display folders and restricted visibility for supporting fields. 4.3 Correct data types and formats The data type tells the model whether a value is a date, number, text or category. The format controls how the value is presented, for example as currency, a percentage or a decimal. An incorrect type can prevent valid grouping and calculations, while an ambiguous format can cause users to misinterpret the result. Data categories should also be assigned where relevant, including for addresses, cities and geographic codes. The model needs to distinguish clearly between order, invoice, shipping and payment dates. A marked date table and explicit measures reduce the number of accidental interpretations. 5. A semantic model that Copilot can understand Microsoft’s tutorial on preparing a semantic model for AI recommends practices that include star schema design, clear naming and reduced complexity. This does not require every model to look the same. The goal is to make relationships, table roles and measure definitions unambiguous. 5.1 Star schema In a star schema, fact tables store events or numeric values and dimension tables provide context such as customer, product, time and region. This structure helps users and AI systems distinguish what is being measured from the dimensions used to break down the result. A complex snowflake model, numerous helper tables and several possible filter paths may be technically justified, but they make interpretation more difficult. If the physical model cannot be simplified, use the AI data schema to narrow the part exposed to Copilot. 5.2 Relationships and filter direction Relationships should reflect the actual data logic. Review their cardinality, active status and filter direction. Bidirectional relationships and multiple alternative paths can produce unexpected results, particularly when a question does not specify enough context. Document inactive relationships and special rules used by selected measures. If analysing sales by shipping date requires different logic from analysing by order date, users need to know how to phrase the question and the model should provide clearly differentiated measures. 5.3 Explicit DAX measures Explicit DAX measures place approved business logic in one location. They provide a safer basis for answers than ad hoc aggregation of numeric columns. Each measure should have a clear name, correct format, useful description and a business owner. Before making a model available to Copilot, check: whether official KPIs are implemented as measures; whether similar or duplicate names remain in the model; whether the result format matches the business meaning; whether measures work correctly across filters and aggregation levels; whether technical fields and helper measures are hidden from users; whether results have been reconciled with reference reports. 6. Prep data for AI in Power BI Prep data for AI is a collection of tools that saves configuration at semantic model level rather than on an individual report. This matters because one model can support multiple reports. A change to the AI schema or instructions may therefore affect more than one use case. Microsoft describes four elements used to prepare a model for natural-language interaction: the AI data schema, verified answers, AI instructions and descriptions. These elements do not work in exactly the same way across every Copilot experience. Test the configuration in the same experiences that users will access. Mechanism Purpose When it is particularly useful What it does not replace AI data schema Defines the subset of tables, columns and measures made available to Copilot When a model is large or contains technical fields and similar measures Data cleansing and correct relationships Verified answers Connect approved visuals to trigger phrases For frequent or ambiguous questions that need a consistent interpretation Source data testing and access controls AI instructions Provide rules, definitions and business context When the organisation uses terms whose meaning cannot be inferred from field names Official measures and an unambiguous model Descriptions Document the meaning of tables, columns and measures When an object’s name does not fully explain its use Instructions that cover rules across the domain 6.1 AI data schema The AI data schema limits the part of the model that Copilot considers when answering questions about data. Do not expose the entire schema by default. Technical tables, key fields, unused measures and objects created only to support a report increase the number of possible interpretations. A well-designed AI schema should include official dimensions, approved measures and the fields required for common analyses. Review the scope with business process owners. A schema that is too narrow will prevent valid questions from being answered, while one that is too broad may increase ambiguity. 6.2 Verified answers Verified answers connect an approved visual to defined trigger phrases. They are useful when users frequently ask about the same indicator or use several terms with a similar meaning. Consider a question about sales by area. In one organisation, area may mean a geographic region; in another, it may refer to a product group. A verified answer can direct the relevant wording to a visual that has already been checked. The underlying measure, filters and permissions still need to be tested. For each verified answer, record an owner, the supported questions, the metric source and the date of the latest review. A change to the indicator definition or report structure should trigger another validation. 6.3 AI instructions AI instructions give Copilot context that the model structure alone cannot express easily. They can explain organisational terminology, preferred measures, rules for interpreting periods and relationships between concepts. For example, an instruction may state that an active customer is one who purchased in the past 90 days and that peak season covers June to August. It should use the exact names of model objects and short, testable rules. Instructions are not a security control and do not guarantee that every rule will be followed. Microsoft notes that the language model treats them as guidance. Important financial and operational definitions should still be implemented through measures, relationships and data governance processes. 6.4 Descriptions for model objects Descriptions should explain an object’s meaning, intended use and relevant limitations. Instead of sales value, specify that the measure represents net revenue after discounts and returns, in the reporting currency and by invoice date. According to Microsoft’s current documentation, descriptions do not affect every Copilot capability in the same way. They are used in selected search and DAX query scenarios. They are still worth maintaining because they improve model documentation and prepare the model for further development of AI features. 7. Business context for Copilot A model can be technically correct and still fail to reflect the language used in the business. The sales team may use the term active customer, finance may refer to recognised revenue and operations may discuss a closed order. Each concept needs a definition and an owner. A business glossary is a useful starting point. It should include: the term and its accepted synonyms; a definition approved by the business owner; the corresponding table, column or measure in Power BI; the applicable time, currency, scope and aggregation rules; exceptions and situations in which the metric should not be used; the person responsible for approving changes. The glossary should not exist only outside Power BI. Transfer its most important information into names, descriptions, measures, AI instructions and verified answers. Otherwise, users and Copilot will continue to work with incomplete context. 8. Data security and governance Copilot makes it easier to ask questions, but the core access principle remains the same: users should only have access to the data required for their role. Review workspace and model permissions, row-level security roles, object-level security, Microsoft Entra groups and the way reports and models are shared. The control design should also reflect Microsoft’s guidance on the privacy, security and responsible use of Copilot in Microsoft Fabric. Write permissions require particular attention. In Power BI, the enforcement of row-level security depends in part on the user’s role and permissions for the model. Testing should use accounts that represent real user roles, not only an author or administrator account. Check whether object names, descriptions and report metadata reveal information that a user should not see. Microsoft’s documentation on using Copilot with semantic models states that, in Power BI Desktop, metadata from the current report page may in some situations be used as grounding data and may contain data values. A security assessment should therefore cover both the records and the descriptive layer of the model. Governance should define: who may prepare a model for AI use; who approves definitions and verified answers; which models may be marked as Approved for Copilot; how often regression tests are run; how incorrect answers are reported and analysed; when a model change requires renewed approval. 9. Testing Copilot answers A test should involve more than one sample question. Build a set of scenarios that reflects the language and needs of users. For each question, define the expected measure, filter scope, source of the reference value and acceptable presentation. Test type Example What to assess Basic question What were net sales last month Selection of the measure, period and value format Synonym Show turnover by region Whether the user’s language maps correctly to the official concept Ambiguous question Show the result for each area Whether Copilot asks for clarification or selects the approved interpretation Complex filters Sales to manufacturing customers in Poland last quarter Accuracy of all filters and relationships Permissions The same question asked by users from different regions Whether each user sees only the permitted data scope Change resilience Repeating tests after a measure or relationship changes Whether the update has degraded previously correct answers Assessment should cover the numeric result and how Copilot arrived at the answer. Where available, diagnostic information about how Copilot created an answer can help with investigation. The explanation of the mechanism is not proof that the result is correct. The approved metric and controlled data set remain the reference point. Copilot outputs are nondeterministic. The same prompt and grounding data can therefore produce different results. In some experiences, however, asking the same question within 24 hours while the model remains unchanged may return a cached answer. Testing is not intended to prove that every answer will always be identical. It should show that the model directs Copilot to the right data, common questions receive correct answers and users understand the known risks and limitations. 10. Common mistakes when preparing Power BI for AI 10.1 Enabling Copilot before cleaning up the model The team focuses on licensing and settings but does not review names, relationships and measures. Copilot is enabled on a model that analysts already found difficult to use. The result is ambiguous answers and a rapid loss of user trust. 10.2 Exposing the entire schema Every table and field is included in the AI data schema, including technical keys, helper measures and unused objects. More elements do not necessarily provide better context. In a large model, they can make it harder to select the correct field. 10.3 Treating AI instructions as a substitute for modelling The instructions attempt to explain dozens of exceptions that should be implemented in measures and model rules. Such a document is difficult to test and maintain. The more important the rule, the stronger the case for enforcing it in the model or data process rather than describing it only in natural language. 10.4 No ownership of metrics Analysts create verified answers without formal confirmation of which indicator definition is authoritative. When results differ, no one knows who can approve a change or which value should be used as the reference. 10.5 Testing only by model authors Model authors know the names and data structure, so they ask questions that fit the design. Business users rely on abbreviations, synonyms and incomplete terms. Tests need to include real questions collected from intended users. 10.6 No regression testing after changes A change to a measure definition, relationship, field name or report can affect earlier scenarios. Without regression testing, the organisation cannot know whether the model remains ready for Copilot. 11. Power BI Copilot readiness checklist [ ] We have identified owners for data, models and key metrics. [ ] Source data is subject to repeatable quality and freshness checks. [ ] Official KPIs have unambiguous definitions and explicit DAX measures. [ ] Table, column and measure names match the language used by the business. [ ] Technical, unused and supporting fields are hidden or removed from the AI scope. [ ] Data types, formats, categories and the date table are configured correctly. [ ] Relationships use the correct cardinality and do not create ambiguous filter paths. [ ] We have defined an AI data schema that includes only the required objects. [ ] Common and ambiguous questions have verified answers where appropriate. [ ] AI instructions explain terminology and rules that cannot be inferred from the model structure. [ ] Tables, columns and measures have useful descriptions. [ ] We have checked tenant settings, capacity, licensing and current feature limitations. [ ] We have tested RLS roles, permissions and model access with test user accounts. [ ] The test set covers basic questions, synonyms, ambiguity and complex filters. [ ] Copilot results are compared with approved reports or reference values. [ ] We have defined a process for reporting incorrect answers and revalidating the model. [ ] If the organisation uses the Approved for Copilot setting, which is currently in preview, the model is marked only after testing is complete. 12. Preparing the organisation to use Copilot Even a well-prepared model will not help if users do not know how to interpret the answers. Implementation should include short training, sample questions, an explanation of the data scope and rules for validating results. Make clear which scenarios provide decision support and which require review by an analyst or process owner. A pilot based on one model with a clearly defined scope is a practical starting point. It allows the team to collect user questions, assess ambiguity and build a test set before extending the feature to other areas. The pilot should have completion criteria, such as correct handling of priority scenarios, approval from metric owners and no critical permission issues. After launch, monitor usage, capacity cost, user reports and changes to Microsoft’s documentation. AI readiness requires ongoing monitoring and another round of testing after material changes. 13. Why TTMS Preparing Power BI for Copilot requires skills in data integration, semantic modelling, Power BI, Microsoft Fabric, security and user adoption. Focusing only on the Copilot interface does not address problems in data sources, transformations and business definitions. An engagement with TTMS can cover an assessment of existing model readiness, improvements to the data layer, changes to relationships and measures, and preparation of the test approach. Prep data for AI configuration should follow confirmation of the environment requirements and agreement on the scope of work. The engagement may focus on one pilot model or a programme spanning multiple domains and teams. A practical engagement may include: An inventory of data sources, models, reports and user groups. An assessment of data quality, model architecture and the risk of ambiguous answers. Agreement on official metrics and a glossary of business concepts. Semantic model optimisation and configuration of the AI data schema, verified answers and AI instructions. Functional, regression, security and performance testing. Preparation of governance rules, documentation and user materials. Model maintenance and renewed validation after changes to data or Microsoft features. The intended result is a model whose structure is clear to analysts and business users, with Copilot answers that can be assessed against approved definitions. The aim is to reduce ambiguity and establish a quality-control process. Even a well-prepared model cannot guarantee error-free generative AI output. 14. Discuss Power BI AI readiness If your organisation uses Power BI and plans to make Copilot available, start with a review of the data and semantic models. TTMS can help define the pilot scope, identify gaps and prepare a roadmap from data sources through to user testing. Contact TTMS to discuss preparing your Power BI and Microsoft Fabric environment for the secure use of AI capabilities. 15. FAQ Who should own a Power BI semantic model prepared for Copilot? A semantic model prepared for Copilot should have clearly assigned owners for data quality, KPI definitions and model maintenance. Ownership helps ensure that business terms, measures and AI configurations remain accurate as data sources, processes and reporting requirements evolve. How large should the AI data schema be? The AI data schema should include only the tables, columns and measures needed for common business questions. Exposing too many technical objects can increase ambiguity, while an overly restrictive schema may prevent Copilot from answering valid questions. Can Power BI Copilot use business terminology that does not exist in the data model? Copilot can better understand business terminology when organisations provide context through measure names, descriptions, AI instructions and verified answers. However, important business concepts should still be represented directly in the semantic model whenever possible. When should a Power BI model be revalidated for Copilot? A model should be revalidated after significant changes to data sources, KPI definitions, relationships, security settings, AI configurations or Microsoft Copilot features. Regular regression testing helps confirm that previously validated scenarios still return correct results. Is preparing a model for Copilot a one-time project? No. Copilot readiness should be treated as an ongoing governance process rather than a one-time implementation task. As data, business definitions and AI capabilities change, organisations need to review model quality, test key scenarios and update documentation on a regular basis.

Read
GPT-6 Astra in Microsoft 365 Copilot: Access, Tasks and Cowork Costs

GPT-6 Astra in Microsoft 365 Copilot: Access, Tasks and Cowork Costs

Does your company use Copilot, and would you like to try GPT-6 Astra? OpenAI’s model is also available in Copilot Cowork. This means you can try it when working with documents, email and calendars in Microsoft’s environment. Access to Astra depends on your organisation’s licences and settings, while tasks performed in Cowork are billed based on credit usage. What does your administrator need to enable? Which tasks can you delegate to Astra in Cowork, and how are they handled in ChatGPT Work? Below, we explain access requirements, differences in working with files and billing rules. For guidance on choosing an assistant for your organisation, see our comparison of Microsoft Copilot and ChatGPT for business. 1. What Does GPT-6 Astra Bring to Copilot Cowork? Microsoft lists GPT-6 Astra among the models available in Copilot Cowork. Users select a model from the list enabled by their organisation. The default Auto setting lets Cowork choose a model for the task; a label next to the response shows which model was used. Selecting Astra applies to work within Cowork. The availability of a particular model in other Copilot features needs to be checked separately. GPT-6 Astra is another model you can assign tasks to in Cowork. Cowork itself provides the tools for finding information, creating files and taking action in Microsoft 365. Your choice of model may affect how information is analysed, the level of detail in the response and the time taken to complete the task. When evaluating Astra, check whether it handles an existing task better: whether it brings together findings more accurately, accounts for exceptions and produces a result that requires fewer revisions. Work IQ gives Cowork access to the context of your organisation’s work. When preparing a project summary, the information needed may be spread across documents, correspondence and meeting materials. Cowork can search for the organisational resources required for the task. Before trying it, check that the employee’s account has access to the relevant materials and that they include the latest decisions and updates. This determines which information Astra will use to produce its result. We discuss the model’s test results and examples of its use in our article GPT-6 Astra: Impressive Achievements and New Possibilities for Business. 2. How Can You Access Astra in Copilot Cowork? For business users, Microsoft describes Cowork as a service that requires a Microsoft 365 Copilot licence and usage-based billing for task execution. An administrator then needs to configure employee access. There are two separate settings to configure: Access to Cowork. The employee must belong to a group covered by a spending policy that includes Cowork. The administrator configures this in the Microsoft 365 admin centre under Copilot, Cost Management, Configuration. This is where they specify the users, budget and billing method. Access to models provided by OpenAI. In the Copilot settings, the administrator specifies which users can use OpenAI as a Microsoft subprocessor. Once you have access, open Cowork and select Astra from the model list. For your first task, check the model label next to the response. If an employee can see Cowork but cannot find Astra, the administrator should check the model provider settings. To make Cowork available to a specific team, the administrator must grant access to the relevant user group. A low credit limit restricts spending while still allowing employees covered by it to get started. 3. Astra in Cowork and ChatGPT Work: Differences in Task Execution Preparing a report involves finding up-to-date data, processing it and saving the result somewhere the team can access. At each stage, the tools available to the model matter. The comparison below shows how the two environments work with the materials needed for a task. Working with Materials in Copilot Cowork and ChatGPT Work Task Component Copilot Cowork ChatGPT Work Finding materials Searches Microsoft 365 resources accessible to the user, including email and files. Plugins can provide access to additional sources. Uses files provided for the task and information retrieved through enabled apps and authorised accounts. Working on documents Creates and modifies documents, spreadsheets and presentations. Output files are saved to the workspace in OneDrive or SharePoint. Creates and edits files. Transferring them to another system depends on the operations supported by the connection to that system. Files stored on the computer A file can be uploaded to the session. Cowork does not edit files directly on the user’s drive. Work in a supported desktop app can use local files once the appropriate access has been granted. Using an application through a browser The local Edge browser uses the employee’s existing sign-in. The feature must be enabled by an administrator. Access depends on the browser tool selected and the permissions granted. A cloud task requires separate authorisation to access company resources. When working through a browser, you need to consider where the task is running. Cowork supports the local browser when its web version is open in Edge. This feature is currently unavailable in the Copilot desktop app and on mobile devices. Cowork and Edge must also use the same work account. If the computer goes to sleep, actions requiring the local browser may be paused. In ChatGPT Work, the model can use shared files and applications on the computer during a local task. A task launched in the cloud runs in a separate environment. If the required materials are stored only on the employee’s drive or are accessible through a company VPN, they need to be made available to that environment through a supported method. As it works, Cowork displays the successive stages of the task. You can interrupt the session, clarify your instructions or provide missing information. Before taking significant actions, such as sending a message or scheduling a meeting, Cowork asks for approval. The additional confirmations it requests also depend on permissions granted earlier. For your first trial, choose a task for which you can clearly identify both the source materials and where the result should be saved. You can find examples of responsibilities in sales, HR, finance and other departments in our overview of 10 practical uses of Microsoft Copilot in an organisation. 4. How Much Does It Cost to Use Astra in Cowork and ChatGPT Work? Your budget needs to cover both the subscription and the use of tools to carry out tasks. For a company that already has the appropriate licences, enabling Cowork primarily means budgeting for usage charges. Below are the public prices for selected business plans. Subscription prices. Charges for task execution are explained below. Plan Monthly Price per User Terms Microsoft 365 Copilot Business EUR 18.20; currently EUR 15.60 under a promotional offer Billed annually, excluding tax. A separate qualifying Microsoft 365 licence is required. Available for up to 300 users. Microsoft 365 Copilot for enterprise EUR 26 Billed annually, excluding tax. A separate qualifying Microsoft 365 licence is required. ChatGPT Business USD 20 when billed annually or USD 25 when billed monthly A minimum of two users. Public pricing in USD; the final amount depends on factors including taxes and the market where the subscription is purchased. ChatGPT Enterprise Custom pricing Usage limits and billing are defined in the agreement. The Copilot Business promotion applies to the first year with an annual commitment and runs from 1 July to 31 December 2026. 4.1 How Are Cowork Tasks Billed? Cowork charges for factors including model usage, context retrieval, tool calls and runtime. Usage is converted into Copilot Credits; under the published pay-as-you-go pricing, one credit costs USD 0.01. One thousand credits therefore cost USD 10. The cost of an individual task depends on the number of credits consumed. The selected reasoning level also affects usage. Cowork offers Light, Medium, High, Extra High and Max settings. A higher level may increase task duration and credit consumption. For recurring work, check whether increasing this setting improves the result enough to justify the cost. 4.2 What Does Credit Usage Mean in ChatGPT Work? ChatGPT Work follows the usage limits and billing rules of the relevant plan. Under agreements based on a shared credit pool, tasks reduce the available balance. Credits already paid for under the agreement are covered by that payment. Additional charges may arise once those credits run out, if the agreement and settings allow work to continue. When comparing costs, use the same set of tasks and output requirements. Record usage, the number of retries and the extent of any revisions needed. Calculating the cost per successfully completed task shows how much you pay for a result your team can use. First, convert each service’s credit usage into a monetary amount using its own pricing. 5. What Data Protection Rules Apply to Astra in Cowork? In Copilot Cowork, Astra is provided by OpenAI as a Microsoft subprocessor. According to the documentation, this use of the model is governed by Microsoft’s terms and Data Protection Addendum, subject to specified exclusions. These services fall within the EU Data Boundary, with documented exceptions. Microsoft currently excludes them from its commitments to process data in a specific country. This detail matters to organisations that require processing exclusively in Poland, for example. When enabling Astra, the administrator should therefore consider the model provider’s policies and access to the materials used in the task. In ChatGPT Work, whether the task runs locally or in the cloud also matters. During a local task, file excerpts, screenshots and tool outputs may be sent to OpenAI. Company AI policies should account for this method of sharing information as well. 6. Which Task Should You Start with When Trying Astra? Start with a responsibility that regularly involves an employee gathering information and preparing material for other people. This workflow lets you assess both Astra’s analysis and the tools available in Cowork or Work. A weekly project summary is one example. A sample prompt for your own trial: Using the project folder [link] and correspondence about this project from the past seven days, prepare a report for the manager. List revised deadlines, pending decisions and the people responsible for next steps. Provide a source and date for each finding. If the materials contain conflicting information, show the discrepancy and explain what you need to resolve it. Save the report as a DOCX file using the attached template in the folder [link]. Draft a message to [recipients] with a link to the report. Leave sending it subject to my approval. Check whether the report reflects the latest decisions and updates, provides sources and dates, identifies conflicting information and assigns responsibilities correctly. Also assess whether it follows the template and whether the file has been saved in a folder accessible to its recipients. After a successful trial, you can consider running the task regularly. Cowork supports scheduled tasks and tasks triggered by events such as an email or a Teams post. By default, event-triggered tasks prepare actions for approval. We discuss how to design the entire process in our guide to business process automation with Copilot. 7. Prepare Your First Astra Tasks with TTMS Through our AI consulting services, we help you determine which data a task requires, which tools need to be made available and how to assess the result. We also analyse the required licences and usage billing arrangements. We combine consulting with AI solution design and the integration of business systems. TTMS was the first company in Poland to obtain accredited ISO/IEC 42001 certification for its artificial intelligence management system. The TÜV Nord Poland audit covered AI design and usage policies, including risk management and project documentation. Tell us which task you would like to delegate to Astra and which applications your team uses. Talk to TTMS about AI consulting for your business. GPT-6 Astra in Copilot Cowork: Frequently Asked Questions Does selecting Astra in Cowork change the model across all Copilot applications? The selection applies to work within Cowork. Microsoft describes a separate model selection option for this environment. To find out which model powers a particular feature in Word, Excel or Teams, check that feature’s documentation. When reviewing a Cowork task, you can see which model was used by checking the label next to the response. Will the same model give an identical response in Copilot and ChatGPT? The result may differ. The model works with the information provided by each product and uses its tools, instructions and reasoning settings. When comparing results, check which materials the model received and which actions it could perform. Only then can you meaningfully assess the differences in the outputs. Why can I see Cowork but cannot select GPT-6 Astra? Access to Cowork and access to OpenAI models are controlled by separate settings. Your administrator should confirm that your account is allowed to use models provided by OpenAI as a Microsoft subprocessor. The model list displayed in Cowork reflects the access granted by your organisation. Does a Microsoft 365 Copilot subscription cover all Cowork tasks? Cowork tasks incur additional usage-based charges. Copilot Credit consumption depends on factors including the model, information retrieval and tools used. Administrators can set spending policies for users and groups. Your budget should account for both the subscription and expected Cowork usage. Can Astra in Cowork edit a document saved on my computer? You can upload a document to a Cowork session. According to the current FAQ, the service does not open or edit files directly on your local drive. Cowork works with the materials you provide and files available in OneDrive and SharePoint. Support for the Edge browser is a separate feature. Will the same Astra model produce the same result in Cowork and ChatGPT Work? The result also depends on the available data, instructions, tools and reasoning settings. A task performed using the same model may therefore proceed differently in the two environments. Comparing results using the same materials will reveal differences in the completeness of the output and the actions performed.

Read
Global Employee Training: 2026 Strategies That Work

Global Employee Training: 2026 Strategies That Work

A sales rep in Manila may need to learn the same product update as an engineer in Warsaw or a compliance officer in Toronto. Yet they work in different languages, time zones, and regulatory environments. For companies running global employee training programs in 2026, this makes a single standardized training deck increasingly impractical. Global training therefore requires a balance between consistency and local relevance. Core knowledge, processes, and brand standards may stay the same across markets, while other parts of the training need to reflect local regulations, language, culture, or job-specific requirements. The challenge is not simply to translate the same course into multiple languages. It is to decide what should remain standardized, what needs to be adapted, and where full localization is necessary. The right approach can make training easier to scale, more relevant for employees, and more consistent across regions. 1. What Global Employee Training Looks Like in 2026 Global employee training in 2026 is increasingly designed around a shared core with room for local adaptation. Companies may standardize product knowledge, internal processes, brand guidelines, or compliance principles, while adjusting language, examples, legal references, and delivery formats for individual markets. Digital learning platforms make this approach easier to manage across regions. Employees in São Paulo and Seoul can complete the same core course while receiving content adapted to their language, role, or local requirements. This helps organizations maintain consistency without if every audience should receive exactly the same version of the training. Artificial intelligence and data analytics have added another layer of sophistication. Training systems now track how individual employees learn, where they struggle, and what content keeps them engaged, then adjust the experience accordingly. Personalization has become the baseline expectation for global employee training and development, not some nice-to-have extra. 1.1 Key Differences from Traditional, Single-Region Training Traditional training is often designed for one language, regulatory environment, and organizational context. Global employee training has to account for several of these at the same time. A compliance or safety course, for example, may need more than a direct translation. Legal terminology, procedures, examples, and even the way instructions are presented can differ between countries. Live training also requires additional planning when employees are spread across time zones. As a result, global training programs are usually built around a combination of standardized and localized content. The key is deciding which elements need to remain consistent across the organization and which should be adapted for a particular market or audience. 1.2 Why This Matters Now: Distributed Teams, AI, and Skills Gaps Distributed teams are the norm rather than the exception, and that alone forces companies to rethink how they train people. Add rapidly evolving AI tools and widening skills gaps across industries, and the pressure to modernize training becomes hard to ignore. Companies that fail to adapt risk losing talent to competitors offering more relevant, more accessible learning experiences. Those that invest in scalable solutions for global employee training are better positioned to keep pace with both technology shifts and workforce expectations. 2. The Business Case: Benefits of Global Employee Training and Development Global employee training supports much more than compliance. It helps companies build the skills they need across different markets, introduce new processes more consistently, reduce operational risk, and give employees a clearer understanding of what is expected of them. Its value is especially visible in organizations that operate across several countries, where differences in skills, regulations, language, and local working practices can quickly create gaps between teams. 2.1 Closing Skills Gaps Across Markets Skills gaps rarely look the same from one region to the next. A well-designed global training program identifies where those gaps exist and builds targeted content to close them, so every market has the competencies it needs to hit business goals. This matters especially as new technologies and processes roll out faster than ever, leaving less room for regional teams to fall behind. 2.2 Boosting Engagement, Confidence, and Retention Worldwide Employees who feel equipped to do their jobs well tend to stick around longer. Strong training programs give people the confidence to take on new responsibilities, and that confidence translates into higher engagement and better retention across every office, not just headquarters. 2.3 Strengthening Compliance and Reducing Regional Risk Regulations differ from country to country, and getting them wrong can be costly. A structured global training approach makes sure compliance training reflects local laws while still aligning with company-wide standards, which keeps costly missteps in any one market to a minimum. 2.4 Building a Consistent Culture Across Borders Culture can fracture quickly across a distributed workforce if there’s no shared thread connecting offices. Training is one of the most effective tools for reinforcing company values and expectations everywhere the business operates. It gives teams a sense of belonging to the same organization, even when they’ve never met face to face. 3. Common Challenges in Training a Global, Distributed Workforce Running training across several countries introduces challenges that are less visible in a single-market program. Language, time zones, local regulations, infrastructure, and differences in learning culture all affect how training should be designed and delivered. The difficulty is usually not creating one course. It is maintaining a program that works across different environments without making it unnecessarily complex or expensive. 3.1 Language and Cultural Diversity Language barriers can quietly undermine even the best-designed course. Beyond translation, cultural context shapes how people read examples, humor, feedback, all of it, and training that ignores this risks losing its audience before the message ever lands. Automated translation tools help with speed, but they still miss idiom and tone often enough that human review remains necessary before content goes live in a new market. 3.2 Time Zone and Logistical Barriers Coordinating live sessions across a dozen time zones is nearly impossible without leaving someone out. That’s pushing companies toward asynchronous, self-paced formats that let employees engage with material on their own schedule instead of forcing everyone into the same time slot. The trade-off: self-paced courses without any live touchpoint or accountability structure tend to see weaker completion rates than blended formats, which is worth weighing before going fully asynchronous. 3.3 Balancing Global Consistency with Local Relevance Lean too hard on standardization and training feels disconnected from local realities. Lean too hard on localization and the company loses a consistent message, plus the cost and coordination burden can outweigh the benefit for smaller or less regulated markets. Striking that balance is one of the harder judgment calls in designing any global employee training and development strategy. 3.4 Technology and Infrastructure Disparities Not every office has the same bandwidth, devices, or digital literacy. Training platforms need to work reliably across varying levels of technological infrastructure, or entire regions risk being left with a worse learning experience than others. 3.5 Measuring Impact Across Multiple Regions Data collected in one market doesn’t always translate cleanly to another. Comparing outcomes across regions requires consistent metrics and reporting tools, otherwise it becomes difficult to know whether the program is working everywhere it’s deployed. 4. FourCore Strategies for Structuring Global Training Programs Companies generally choose from four broad approaches when structuring global training, each with its own trade-offs between simplicity and personalization. Strategy 1: Fully Standardized Training for All Topics This approach delivers identical content everywhere. It’s the easiest to build and maintain, but it risks missing the cultural and regulatory details that matter in specific markets. Strategy 2: Standardized Approach, Customized by Topic Here, some topics stay uniform across the company while others get adapted per region. This gives more flexibility than a fully standardized model without the resource demands of full localization. Strategy 3: Shared Objectives with Region-Specific Content Under this strategy, every region works toward the same learning objectives but builds content that fits local context. It’s a middle ground that keeps the company aligned while respecting regional differences. Strategy 4: Fully Localized Objectives and Content per Region This is the most tailored approach, with both objectives and content built specifically for each market. It delivers the most relevant experience but demands significant time, budget, and coordination, and it’s often overkilled for smaller regional offices or lightly regulated topics where a shared, lightly adapted version works just as well. Choosing the Right Strategy for Your Organization The right strategy depends on company size, industry regulation, and how much variation exists between regional teams. Organizations with tighter budgets often start with a standardized core and layer in customization as they scale, while larger multinationals with complex regulatory needs may need full localization from day one. 5. Building Blocks of an Effective Global Training and Development Program A global training strategy needs to translate into a program that employees can actually use across different countries, roles, and working environments. That means deciding what people need to learn, which content should be shared globally, where local adaptation is necessary, and how employees will access the training. 5.1 Types of Training to Include: Technical, Compliance, Leadership, and Soft Skills Most global training programs cover several different areas. These may include role-specific technical skills, compliance and safety training, leadership development, product knowledge, and soft skills such as communication or teamwork. They do not all require the same approach. Product or process training can often use a common global core, while compliance content may need significant changes to reflect local regulations. Leadership and communication training may also need different examples or scenarios depending on the cultural and organizational context. 5.2 Localization and Multilingual Content Delivery Localization goes beyond swapping words from one language to another. It means adjusting examples, tone, and even visual design so the material feels natural to the audience. Multilingual delivery has become a baseline expectation for any employee training platform for global companies serving a diverse workforce. 5.3 Blended and Self-Paced Learning Models for Different Time Zones Combining live sessions with self-paced modules gives employees flexibility to learn when it suits them, without losing the benefits of interactive discussion when schedules do align. This blended model has become one of the most practical answers to the time zone problem. 5.4 Peer Learning and Regional Mentorship Networks Some knowledge is easier to develop through interaction with colleagues than through a course alone. Regional mentors, subject-matter experts, and peer groups can help employees apply what they have learned to their actual work. They can also answer questions that are specific to a particular market, customer group, or local process. This is especially useful after formal training has finished, when employees start applying new knowledge in day-to-day situations. 6. Using AI and Technology to Scale Training for Skills Development Technology makes it possible to deliver and manage training across large, distributed teams. AI can support this process by helping organizations personalize learning, adapt content, translate materials, and analyze training data. Human review is still important, especially when content involves compliance, safety, culture, or sensitive terminology. 6.1 AI-Driven Personalization and Adaptive Learning Paths AI can adjust a learning path in real time based on how an individual employee is progressing. It gives someone more practice where they’re struggling and pushes them faster through material they’ve already nailed down. This kind of personalization would be nearly impossible to manage manually across a large, distributed workforce. 6.2 Automated Translation and Localization Tools Automated translation tools speed up the process of adapting content for multiple markets, cutting both cost and turnaround time. Paired with human review for cultural accuracy, these tools make multilingual delivery far more manageable than it used to be, though relying on machine translation alone still creates a real risk of tone-deaf or awkward phrasing in markets with limited review. 6.3 Learning Analytics for Real-Time Performance Insights Learning analytics help L&D teams understand how employees are progressing across courses and regions. They can show completion rates, assessment results, engagement with individual modules, or areas where learners repeatedly encounter difficulties. This data can be used to improve existing courses, identify skills gaps, and decide where additional training or support is needed. It also gives global training teams a more consistent way to compare results across markets. TTMS supports organizations in building and maintaining this type of learning environment through its AI Solutions and E-Learning administration services. Depending on the scale and complexity of the program, this may include AI-assisted content creation and adaptation, multilingual course management, hosting, reporting, and ongoing updates. The level of technology should match the actual training needs. A large international program may benefit from automation and advanced analytics, while a smaller rollout can often be managed effectively with a simpler platform and a well-defined review process. 7. How to Implement a Global Training Strategy Rolling out a global training strategy starts with a clear assessment of organizational needs and a definition of what success should look like. From there, companies select the training methods and delivery formats that fit their workforce, whether that means blended learning, mobile-first content, or live regional workshops. Engaging local stakeholders early is essential, since they’re the ones who know which cultural or regulatory details need attention before content goes live. Technology plays a central role in execution. An employee training platform for global companies needs to handle content hosting, multilingual delivery, and progress tracking, ideally within a single system rather than a patchwork of tools. Continuous evaluation and feedback loops then let teams refine the program over time, rather than treating the initial rollout as a finished product. TTMS can also support the operational side of a global training rollout by automating processes around course assignment, approvals, reminders, and completion tracking. With Process Automation and Low-Code Power Apps, these workflows can be connected across departments and regional offices, reducing the need to manage them through separate spreadsheets or manual email exchanges. For organizations already using Microsoft 365 and Azure, training processes can also be integrated with the tools employees and administrators use every day. 8. Measuring ROI and Impact of Corporate Training Programs Globally Proving the value of a global training investment requires looking at more than completion rates. Organizations should track performance indicators tied to productivity, retention, and skill application on the job, alongside qualitative feedback that reveals how employees perceive the training’s usefulness. Comparing outcomes between trained and untrained groups offers one of the clearest ways to demonstrate tangible impact and gives leadership the evidence it needs to justify continued investment or adjust course where results fall short. Business Intelligence tools, such as Snowflake DWH and Power BI, can play a useful role here by consolidating training data from multiple regions into a single view. That makes it far easier to spot trends and report results across the organization instead of reviewing each market’s numbers in isolation. 9. Real-World Examples of Global Employee Training Done Right Successful global employee training starts with matching the learning format to the content, audience, and business context. Some topics can be delivered through standardized materials across regions, while others require a more tailored approach because of local regulations, safety requirements, language, or cultural differences. A good example comes from a global production and technology company that needed to standardize Health & Safety training for production and office employees across five locations worldwide. Previously, individual sites used different materials, which made it difficult to ensure that employees received the same information and that training completion was properly tracked. TTMS developed a single interactive e-learning course built around workplace scenarios and storytelling. Employees worked through situations that could lead to accidents and selected the appropriate response, receiving immediate feedback on their decisions. The course helped the company deliver the same core safety principles across a multicultural workforce while replacing part of its previously time-consuming classroom training. The platform also gave managers visibility into who had started or completed the training and automatically reminded employees about approaching deadlines. According to the case study, the organization subsequently recorded fewer accidents across its locations. This example shows an important principle of global employee training: not every subject should be handled in the same way. Compliance and safety content often needs more careful adaptation and stronger learner engagement, while other training can remain more standardized. Technology makes it easier to distribute and update learning across locations, but the format and level of localization should still reflect the needs of each audience. If you are planning to scale employee training across countries, languages, or business units, TTMS can help you choose the right mix of standardization, localization, technology, and content formats. Talk to our e-learning experts about your training needs and the best way to structure a global program. Frequently Asked Questions What is global employee training? Global employee training refers to the systematic development of skills and knowledge among employees across different regions and cultures, ensuring that training is relevant, accessible, and effective for a diverse workforce. How to manage global employee training? Managing global employee training involves understanding cultural differences, using technology to improve accessibility, and making sure training content is both standardized and localized to meet regional needs. How do companies handle language barriers in global training? Companies address language barriers by localizing training content, using multilingual support, and employing translation tools to ensure that all employees can understand and engage with the training material. What’s the difference between standardized and localized training? Standardized training provides a uniform approach across all regions, while localized training adapts content to fit the specific cultural and linguistic needs of different employee groups. How do you measure the success of a global training program? Success can be measured through various metrics, including employee performance improvements, retention rates, engagement levels, and feedback from participants regarding the training’s relevance and effectiveness. Building an effective global employee training program takes more than good intentions. It needs the right mix of strategy, localization, and technology, plus a partner who knows how to bring those pieces together. Companies exploring global employee training management software or looking to modernize their approach to workforce learning can turn to TTMS for guidance grounded in real IT implementation experience across AI, automation, and e-learning administration.

Read
1…345…50

The world’s largest corporations have trusted us

Wiktor Janicki

We hereby declare that Transition Technologies MS provides IT services on time, with high quality and in accordance with the signed agreement. We recommend TTMS as a trustworthy and reliable provider of Salesforce IT services.

Read more
Julien Guillot Schneider Electric

TTMS has really helped us thorough the years in the field of configuration and management of protection relays with the use of various technologies. I do confirm, that the services provided by TTMS are implemented in a timely manner, in accordance with the agreement and duly.

Read more

Ready to take your business to the next level?

Let’s talk about how TTMS can help.

TTMC Contact person
Monika Radomska

Sales Manager