GPT-6.1 Sol for business: Greater capabilities, lower costs

Table of contents

    On 29 September 2026, at DevDay, OpenAI introduced GPT-6.1 Sol, just one week after the launch of GPT-6 Sol. The company aimed to bring Sol closer to GPT-6 Astra’s capabilities in coding, document analysis and application use, while keeping the model’s existing base prices.
    OpenAI kept GPT-6 Sol’s pricing: USD 2 per million input tokens and USD 10 per million output tokens, one fifth of Astra’s standard rates. OpenAI also links the update to the development of AI agents, programs that complete tasks step by step using available tools. Analysing documentation, preparing a code change and checking that it works may require multiple calls to the model. The cost of this work grows with the number of operations. GPT-6.1 Sol is intended to let application developers tackle more demanding tasks within the same budget.

    GPT-6.1 Sol

    1. What is GPT-6.1 Sol, and what has changed since GPT-6 Sol?

    GPT-6.1 Sol builds on GPT-6 Sol and is designed for complex tasks involving code, documents and applications. OpenAI’s documentation also lists support for tools, including search and code execution. The model can be used in an AI agent connected to a company’s data sources and tools. An agent could read a customer enquiry, find the relevant specification, compare requirements and draft a response. The company must define which data the agent can access, which actions it can perform independently and which require an employee’s approval.

    Base prices for input and output tokens remain the same as for GPT-6 Sol. The rate for reusing cached data, however, falls from USD 0.20 to USD 0.10 per million tokens. Companies using GPT-6 Sol should test both versions on the same tasks and compare the quality of their results, the number of retries and token usage.

    2. What do benchmark results tell us about business applications?

    OpenAI publishes benchmark results covering document analysis and multistep tasks. When comparing results, check the reasoning level used. A higher level allows the model to spend more computation on analysing a task, which may increase token usage.

    Benchmark Area assessed Result reported by OpenAI
    GDP.pdf Answering questions based on complex PDF documents. Results close to Astra’s at approximately one fifth of the cost per task.
    AutomationBench 1.0.6 Multistep tasks in business applications. A lead of 4.8 percentage points over GPT-6 Sol at the medium reasoning level.
    OSWorld 2.0 Using desktop applications. A score 2.1 percentage points below Astra’s. The cost per task is approximately one seventh of Astra’s. Both models were tested at the maximum reasoning level.

    Each benchmark assesses different model capabilities, so their results should be considered separately.

    • GDP.pdf, developed by Surge AI, tests whether a model can analyse workplace documents and answer questions about their contents. The model must also read tables, diagrams and footnotes correctly and connect information from different parts of a document. The result therefore shows how it handles materials similar to those reviewed by procurement departments or technical teams.
    • Zapier’s AutomationBench assesses task execution in simulated applications. It checks whether the model actually made the required changes, such as updating the correct record and completing every requested action. This matters for businesses: a model’s claim that it has completed a task should be supported by data saved in the system.
    • OSWorld tests computer use, including actions performed through application interfaces. The result cited by OpenAI includes credit for partial task completion. A high score may therefore mean that the model completed most steps correctly but missed the final one, such as submitting a form. When assessing a model’s suitability for automation, companies should also check how often it completes the entire task.

    These benchmark results help companies choose use cases for a pilot. Testing the model on their own documents and in the applications they use will show whether it meets their requirements and how much time is needed to check and correct its work.

    3. How much does GPT-6.1 Sol cost, and how should you read the pricing table?

    API usage, which allows an application to access the model, is billed in tokens: units of processed text that often correspond to parts of words. Input includes instructions and supplied materials. Billable output tokens also include reasoning tokens, which are not visible in the response.

    Rates in USD per million tokens in Standard mode, for requests containing up to 272,000 input tokens:

    Model Input Cache read Cache write Output
    GPT-6 Sol 2.00 0.20 2.50 10.00
    GPT-6.1 Sol 2.00 0.10 2.50 10.00
    GPT-6 Astra 10.00 1.00 12.50 50.00

    This comparison covers standard processing. Faster processing modes and requests with more input data have separate rates, and processing in a selected region may incur a surcharge. Caching can reduce the cost of subsequent requests that use the same instructions or documents. Savings depend on how much data can be reused. The total cost of a task includes token usage, caching charges and the tools used.

    4. Example: how much could you save on 1,000 tasks per month?

    Let’s work through an example: a company completes 1,000 tasks per month, each using 20,000 input tokens and 3,000 billable output tokens, including reasoning. We assume identical usage for both models. To keep the calculation clear, we exclude cache reads and writes, tool fees and surcharges. Token usage must be measured in an actual deployment.

    Monthly cost component GPT-6.1 Sol GPT-6 Astra
    20 million input tokens USD 40 USD 200
    3 million output tokens USD 30 USD 150
    Total token cost USD 70 USD 350

    Token savings amount to USD 280 per month. Now let’s add another assumption: an hour of work by the person checking the results costs the company USD 30. If checking Sol’s results took an average of one extra minute per task, monthly labour costs would rise by USD 500. In this example, using Sol would cost the company USD 220 more per month overall than using Astra. If checking Sol’s results takes the same amount of time as checking Astra’s, or less, lower token spending may reduce the total cost of completing a task. The pilot should compare result quality, model costs and the time spent checking outputs.

    To calculate the cost per successfully completed task, divide the total process costs by the number of tasks completed according to the agreed requirements. The calculation should also include failed attempts, corrections, exception handling and maintenance. Data preparation and system integration costs should be spread over the planned period of use.

    GPT-6.1 Sol cost savings

    5. Which GPT-6.1 Sol use cases are worth testing in your company?

    A pilot could cover comparing documents and proposals, preparing information for customer service and developing code fixes.

    5.1 Comparing technical documents and proposals

    A procurement department receives proposals containing specifications, parameter tables and service terms. The system could prepare a comparison against an agreed set of requirements. For each value, it should identify the source document and the location where it found the information, and flag any missing data. The test set should include two versions of the same specification, parameters expressed in different units and a qualification stated in a footnote. The reviewer can then check whether the system selected the current version and interpreted the qualifications correctly. The total time needed to prepare and check the comparison should be compared with the time needed to complete the task without AI.

    5.2 Preparing customer service information across multiple systems

    After receiving a sales enquiry, an agent could find the customer’s account history, check the status of previous support tickets and prepare an entry in a customer relationship management (CRM) system. This requires connections to the relevant applications and access to the necessary data. The first stage of deployment could end with proposed changes being presented to an employee. Once approved, the system saves them and checks the result. Tests should include records for different customers with similar names to check whether the agent selects the correct one. The same task should also be assigned again to check whether the agent recognises the existing entry and avoids creating a duplicate.

    5.3 Diagnosing bugs and preparing software changes

    An IT team can compare the models when analysing tickets, reproducing bugs and preparing fixes with tests. Previously resolved tickets, where the cause of the bug and the required software behaviour are known, make useful test cases. The assessment should cover the correctness of the change, test results and the time needed for a developer’s review. It should also check whether the model limited its modifications to the necessary scope. Too many code changes can make the review longer and make it harder to identify which changes fix the reported bug.

    6. How can you compare GPT-6.1 Sol with Astra in your own process?

    OpenAI’s model selection guidance recommends comparing Sol and Astra on the same task. The comparison should cover result quality, completion time and task cost under conditions close to the intended use.

      1. Choose a task with a clear completion criterion. This could be comparing all required proposal parameters or preparing a code change that passes specified tests.
      2. Build a set of examples from day-to-day work. Include typical cases, incomplete data and known exceptions. Keep the materials used to refine instructions separate from those used to evaluate the final configuration.
      3. Ensure comparable conditions. Give the models the same materials, tools and permissions. Record reasoning settings and limits. You can then adjust each model’s settings to meet the same quality requirements.
      4. Calculate the full cost and check the result. Measure correctness, retries, system processing time and human review time. Also record cases that require manual completion.
      5. Expand use once the agreed conditions are met. Define which results can be used immediately, which require approval and when a case should be referred to a specialist.

    OpenAI recommends testing AI systems on tasks that reflect day-to-day work and having their results assessed by people with relevant expertise. Tests should be repeated after changes to the model, instructions or data sources. You can also test a division of work between the models. Sol would handle cases where it meets the required quality standard, while Astra would handle selected, more difficult cases. Define in advance when a task should be passed to Astra, for example when Sol misinterprets documents or produces an incorrect response. If required data is missing, it should be supplied before another analysis. The cost of that additional analysis must be included in the calculation. The final division of work will depend on the test results.

    7. Let’s discuss automation in your company

    At TTMS, we help select AI solutions for specific processes and integrate them with business systems. If you are considering GPT-6.1 Sol, contact us. Together, we can identify tasks worth including in a pilot and determine how to assess result quality and the full cost of completing them.

     

    Can GPT-6.1 Sol be connected to a company knowledge base?

    Yes. An application can provide the model with information retrieved from approved company materials. This can be done using retrieval mechanisms described in OpenAI’s documentation. The approach in which a system first retrieves relevant source passages and then uses them to prepare an answer is known as retrieval-augmented generation (RAG). Implementation requires keeping documents up to date, identifying versions and respecting user access permissions. Providing sources helps users check the answers. Performance should also be assessed on questions for which the knowledge base does not contain enough information.

    Does switching from GPT-6 Sol to GPT-6.1 Sol require changes to the application?

    The scope of changes depends on how the application is integrated. The team should check the available interfaces, reasoning settings, tool support and response formats. For example, the GPT-6.1 Sol model documentation states that tool use requires the Responses API and that the none reasoning setting is not supported. This may affect applications built for the previous version. After adjusting the configuration, the team should test it against a saved set of tasks, including error handling. It is worth retaining the option to revert to the previous configuration if the new one does not meet requirements.

    What is GPT-6.1 Sol?

    GPT-6.1 Sol is an OpenAI model designed for tasks involving code, documents and business applications. It builds on GPT-6 Sol while keeping the same base prices for input and output tokens. It also supports tools, allowing developers to use it in AI agents that retrieve information and perform actions across connected systems. For businesses, its appeal lies in the balance between capability and cost. Its suitability for a particular process should be assessed through tests that measure result quality, token usage and the time needed for human review.

    Wiktor Janicki

    We hereby declare that Transition Technologies MS provides IT services on time, with high quality and in accordance with the signed agreement. We recommend TTMS as a trustworthy and reliable provider of Salesforce IT services.

    Read more
    Julien Guillot Schneider Electric

    TTMS has really helped us thorough the years in the field of configuration and management of protection relays with the use of various technologies. I do confirm, that the services provided by TTMS are implemented in a timely manner, in accordance with the agreement and duly.

    Read more

    Ready to take your business to the next level?

    Let’s talk about how TTMS can help.

    Monika Radomska

    Sales Manager