OpenAI

GPT-3.5 Deprecated

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

May 28, 2023 16.4K context 4,000 tokens output
Text Tools Structured Output

Model Overview

High-signal model metadata in a structured two-column overview table.

Provider

The entity that provides this model.

OpenAI

Model ID

The routed model identifier exposed by upstream providers.

openai/gpt-3.5-turbo

Input Context Window

The number of tokens supported by the input context window.

16.4K tokens

Maximum Output Tokens

The number of tokens that can be generated by the model in a single request.

4,000 tokens tokens

Open Source

Whether the model's code is available for public use.

No

Release Date

When the model was first released.

May 28, 2023 3 years ago

Knowledge Cut-off Date

When the model's knowledge was last updated.

2021-09-30

API Providers

The providers that offer this model. This is not an exhaustive list.

OpenAI

Modalities

Types of data this model can process.

Text

What is GPT-3.5 Deprecated

A fuller summary of positioning, capabilities, and source-specific details for GPT-3.5 Deprecated.

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks.

Training data up to Sep 2021.

Capabilities

What GPT-3.5 Deprecated supports

JSON

Structured Outputs

Structured output settings are exposed through OpenRouter for schema-driven or format-controlled responses.

TL

Tool Calling

Tool invocation and tool selection are supported in the routed OpenRouter interface for this model.

MM

Multimodal I/O

This model accepts text input and returns text output.

CTX

Large Context Window

OpenRouter currently lists a context window of 16.4K with up to 4,000 tokens maximum output tokens.

Pricing for GPT-3.5 Deprecated

Primary API pricing shown in the same “quick compare” spirit as the reference page.

Price Comparison

Additional usage-cost dimensions synced into the project for this model.

maxTemperature 2
maxResponseSize 4,000 tokens

API Access & Providers

Places where this model is available, based on the synced detail-page metadata.

OpenAI

Provider Endpoints

Endpoint-level provider data currently available for this model.

OpenAI

Max output: 4,096 1d uptime: 100.0% Supported params: 14 Implicit caching: No

Resources & Documentation

Official model cards, release notes, docs, and other references synced from the source page.

Related Daily Briefs

Recent daily stories tied to GPT-3.5 Deprecated through direct model mentions or provider-level coverage.

Community discussion

What people think about GPT-3.5 Deprecated

GPT-3.5 Deprecated discussions are most active in r/OpenAI, r/LocalLLaMA, r/ChatGPT. Top Reddit threads cluster around benchmark and model-comparison threads, coding workflow discussions.

The strongest match in this snapshot has 262 upvotes and 112 comments.

[This Twitter thread](https://twitter.com/GrantSlatton/status/1703913578036904431) ([Nitter alternative](https://nitter.net/GrantSlatton/status/1703913578036904431) for those who aren't logged into Twitter and want to see the full thread) claims that [OpenAI's new language model gpt-3.5-turbo-instruct](https://analyticsindiamag.com/openai-releases-gpt-3-5-turbo-instruct/) can "readily" beat Lichess Stockfish level 4 ([Lichess Stockfish level and its rating](https://lichess.org/@/MagoGG/blog/stockfish-level-and-its-rating/CvL5k0jL)) and has a chess rating of "around 1800 Elo." [This tweet](https://twitter.com/nabeelqu/status/1703961405999759638) shows the style of prompts that are being used to get these results with the new language model.

I used website parrotchess\[dot\]com (discovered [here](https://twitter.com/OwariDa/status/1704179448013070560)) (EDIT: parrotchess doesn't exist anymore, as of March 7, 2024) to play multiple games of chess purportedly pitting this new language model vs. various levels at website Lichess, which supposedly uses Fairy-Stockfish 14 according to the Lichess user interface. My current results for all completed games: The language model is 5-0 vs. Fairy-Stockfish 14 level 5 ([game 1](https://lichess.org/eGSWJtNq), [game 2](https://lichess.org/pN7K9bdS), [game 3](https://lichess.org/aK4jQvdo), [game 4](https://lichess.org/S9SGg8YI), [game 5](https://lichess.org/OqzdkDhE)), and 2-5 vs. Fairy-Stockfish 14 level 6 ([game 1](https://lichess.org/zP68C6H4), [game 2](https://lichess.org/4XKUIDh1), [game 3](https://lichess.org/1zTasRRp), [game 4](https://lichess.org/lH1EMqJQ), [game 5](https://lichess.org/mdFlTbMn), [game 6](https://lichess.org/HqmELNhw), [game 7](https://lichess.org/inWVs05Q)). Not included in the tally are games that I had to abort because the parrotchess user interface stalled (5 instances), because I accidentally copied a move incorrectly in the parrotchess user interface (numerous instances), or because the parrotchess user interface doesn't allow the promotion of a pawn to anything other than queen (1 instance). **Update: There could have been up to 5 additional losses - the number of times the parrotchess user interface stalled - that would have been recorded in this tally if** [this language model resignation bug](https://twitter.com/OwariDa/status/1705894692603269503) **hadn't been present. Also, the quality of play of some online chess bots can perhaps vary depending on the speed of the user's hardware.**

The following is a screenshot from parrotchess showing the end state of the first game vs. Fairy-Stockfish 14 level 5:

https://preview.redd.it/4ahi32xgjmpb1.jpg?width=432&format=pjpg&auto=webp&s=7fbb68371ca4257bed15ab2828fab58047f194a4

The game results in this paragraph are from using parrotchess after the forementioned resignation bug was fixed. The language model is 0-1 vs. Fairy-Stockfish level 7 ([game 1](https://lichess.org/Se3t7syX)), and 0-1 vs. Fairy-Stockfish 14 level 8 ([game 1](https://lichess.org/j3W2OwrP)).

There is [one known scenario](https://twitter.com/OwariDa/status/1706823943305167077) ([Nitter alternative](https://nitter.net/OwariDa/status/1706823943305167077)) in which the new language model purportedly generated an illegal move using language model sampling temperature of 0. Previous purported illegal moves that the parrotchess developer examined [turned out](https://twitter.com/OwariDa/status/1706765203130515642) ([Nitter alternative](https://nitter.net/OwariDa/status/1706765203130515642)) to be due to parrotchess bugs.

There are several other ways to play chess against the new language model if you have access to the OpenAI API. The first way is to use the OpenAI Playground as shown in [this video](https://www.youtube.com/watch?v=CReHXhmMprg). The second way is chess web app gptchess\[dot\]vercel\[dot\]app (discovered in [this Twitter thread](https://twitter.com/willdepue/status/1703974001717154191) / [Nitter thread](https://nitter.net/willdepue/status/1703974001717154191)). Third, another person modified that chess web app to additionally allow various levels of the Stockfish chess engine to autoplay, resulting in chess web app chessgpt-stockfish\[dot\]vercel\[dot\]app (discovered in [this tweet](https://twitter.com/paul_cal/status/1704466755110793455)).

Results from other people:

a) Results from hundreds of games in blog post [Debunking the Chessboard: Confronting GPTs Against Chess Engines to Estimate Elo Ratings and Assess Legal Move Abilities](https://blog.mathieuacher.com/GPTsChessEloRatingLegalMoves/).

b) Results from 150 games: [GPT-3.5-instruct beats GPT-4 at chess and is a \~1800 ELO chess player. Results of 150 games of GPT-3.5 vs stockfish and 30 of GPT-3.5 vs GPT-4](https://www.reddit.com/r/MachineLearning/comments/16q81fh/d_gpt35instruct_beats_gpt4_at_chess_and_is_a_1800/). [Post #2](https://www.reddit.com/r/chess/comments/16q8a3b/new_openai_model_gpt35instruct_is_a_1800_elo/). The developer later noted that due to bugs the legal move rate [was](https://twitter.com/a_karvonen/status/1706057268305809632) actually above 99.9%. It should also be noted that these results [didn't use](https://www.reddit.com/r/chess/comments/16q8a3b/comment/k1wgg0j/) a language model sampling temperature of 0, which I believe could have induced illegal moves.

c) Chess bot [gpt35-turbo-instruct](https://lichess.org/@/gpt35-turbo-instruct/all) at website Lichess.

d) Chess bot [konaz](https://lichess.org/@/konaz/all) at website Lichess.

From blog post [Playing chess with large language models](https://nicholas.carlini.com/writing/2023/chess-llm.html):

>Computers have been better than humans at chess for at least the last 25 years. And for the past five years, deep learning models have been better than the best humans. But until this week, in order to be good at chess, a machine learning model had to be explicitly designed to play games: it had to be told explicitly that there was an 8x8 board, that there were different pieces, how each of them moved, and what the goal of the game was. Then it had to be trained with reinforcement learning agaist itself. And then it would win.
>
>This all changed on Monday, when OpenAI released GPT-3.5-turbo-instruct, an instruction-tuned language model that was designed to just write English text, but that people on the internet quickly discovered can play chess at, roughly, the level of skilled human players.

Post [Chess as a case study in hidden capabilities in ChatGPT](https://www.lesswrong.com/posts/F6vH6fr8ngo7csDdf/chess-as-a-case-study-in-hidden-capabilities-in-chatgpt) from last month covers a different prompting style used for the older chat-based GPT 3.5 Turbo language model. If I recall correctly from my tests with ChatGPT-3.5, using that prompt style with the older language model can defeat Stockfish level 2 at Lichess, but I haven't been successful in using it to beat Stockfish level 3. In my tests, both the quality of play and frequency of illegal attempted moves seems to be better with the new prompt style with the new language model compared to the older prompt style with the older language model.

Related article: [Large Language Model: world models or surface statistics?](https://thegradient.pub/othello/)

P.S. Since some people claim that language model gpt-3.5-turbo-instruct is always playing moves memorized from the training dataset, I searched for data on the uniqueness of chess positions. From [this video](https://youtu.be/DpXy041BIlA?t=2225), we see that for a certain game dataset there were 763,331,945 chess positions encountered in an unknown number of games without removing duplicate chess positions, 597,725,848 different chess positions reached, and 582,337,984 different chess positions that were reached only once. Therefore, for that game dataset the probability that a chess position in a game was reached only once is 582337984 / 763331945 = 76.3%. For the larger dataset [cited](https://youtu.be/DpXy041BIlA?t=2187) in that video, there are approximately (506,000,000 - 200,000) games in the dataset (per [this paper](http://tom7.org/chess/survival.pdf)), and 21,553,382,902 different game positions encountered. Each game in the larger dataset added a mean of approximately 21,553,382,902 / (506,000,000 - 200,000) = 42.6 different chess positions to the dataset. For [this different dataset](https://lichess.org/blog/Vs0xMTAAAD4We4Ey/opening-explorer) of \~12 million games, \~390 million different chess positions were encountered. Each game in this different dataset added a mean of approximately (390 million / 12 million) = 32.5 different chess positions to the dataset. From the aforementioned numbers, we can conclude that a strategy of playing only moves memorized from a game dataset would fare poorly because there are not rarely new chess games that have chess positions that are not present in the game dataset.

Open Reddit thread

I don't think many people know about Mistral Medium but this model is really good! Mistral website says it's based on an internal prototype:

https://docs.mistral.ai/platform/endpoints/

Even on my Mac Studio with M2 Ultra, it takes a long time to even load the Mixtral 8x7b model (gguf version). That's fine with me; not every model is supposed to run locally. But as long as there's real competition for OpenAI, I think this space will only get better!

Open Reddit thread
r/LocalLLaMA 176 upvotes 96 comments April 4, 2024
GPT-3.5-Turbo is most likely the same size as Mixtral-8x7B

**EDIT**: This prediction has changed in light of new calculations brought to my attention from u/[NighthawkT42](https://www.reddit.com/user/NighthawkT42/). More information at the bottom of the post. I left the main post the same except for the update section which gives a much better estimate for GPT-3.5-Turbo.

[This is a continuation of my last post](https://www.reddit.com/r/LocalLLaMA/comments/1btpk4h/logits_of_apiprotected_llms_leak_proprietary/), but in summary, the paper authors of "[Logits of API-Protected LLMs Leak Proprietary Information](https://arxiv.org/abs/2403.09539v2)" describe how they figured out and exploited a "softmax bottleneck" when calling on an API-Protected LLM over a ton of API calls, which they then used to get a close estimate that GPT-3.5-Turbo's embedding size of around 4096 ± 512. They then mention how this makes GPT-3.5-Turbo either a 7B dense model (by looking at other models with a known embedding size of \~4096), or a MoE that is Nx7B (this has changed and most likely not true, see update section).

Since my last post I have done some thinking and I make the prediction that GPT-3.5-Turbo (the one that has been used since early 2023, not the original GPT-3.5) is most likely around a 8x7B model (this has changed and most likely not true, see update section).

Evidence points to this too indirectly when we take a look at Mixtral-8x7B. Mixtral-8x7B has used by many with the general consensus of this model being on-par or slightly exceeding GPT-3.5-Turbo on most things.

[GPT-3.5-Turbo-0613 & Mixtral-8x7B-Instruct-v0.1 on the LMSYS Chatbot Arena Leaderboard having an averaged ELO difference of \~1 point, though there could be a deviation with GPT-3.5-Turbo-0613 by either +3 or -4 points.](https://preview.redd.it/leg1x5zjxcsc1.png?width=1080&format=png&auto=webp&s=14444a2652db1486db15dbae6865796eb9c4097c)

While the evidence points to this, some still might not think that GPT-3.5-Turbo is around the same size as Mixtral-8x7B because of difference in other language performance, but this could be due to differences in training data. We have no idea what training data was used for both Mixtral-8x7B or GPT-3.5-Turbo, so differences in their performance in relation to training data can be because of that.

Differences can also be found in the tuning of these two models of course as GPT-3.5-Turbo is fine-tuned on RLHF data that has a LOT of human feedback by including the feature for people to vote on an answer that the LLM gives out (ChatGPT), while Mixtral-8x7B-Instruct is a more general Instruction fine-tune.

The use of a MoE by OpenAI makes a lot of sense too. They originally released GPT-3.5 back in November of 2022 inside of ChatGPT, which they thought not many people would use, so compute was not much of a concern. When ChatGPT blew up in the next two months, compute was now the MAIN concern as 10s of millions of people were now using ChatGPT and that model, with more and more people jumping on it in the coming months. They needed a new model that can be as smart as the original GPT-3.5, but able to be served to millions of people at the same time to keep up with very heavy demand.

OpenAI had just finish training GPT-4 not too long ago, which used a 8x MoE (based on indirect knowledge), it showed great promise for the power it can give but also for efficiency in running it compared to a fully dense 1T+ model. They possibly figured that a smaller MoE could possibly get similar performance to GPT-3.5 while costing much less compute to serve to many people, only needing the VRAM to load the full model into GPU memory. If we assume they loaded 2 experts out of a possible 8 experts for serving to people (similar to default Mixtral-8x7B), this would basically quadruple their existing computing power to serve to the growing user base of ChatGPT.

They first released GPT-3.5-Turbo in the API and in ChatGPT Plus to get a better idea of the performance of the model from the public compared to the original GPT-3.5, as well as set aside enough compute to run this model at full ChatGPT scale, which they then did a little while later. They possibly used RLHF data from the public as well to tune the new GPT-3.5-Turbo model to act a lot like ChatGPT as well in most cases, which helped them seamlessly change the model in ChatGPT without most users noticing directly.

Based on ALL of this, I can very confidently say that I think that GPT-3.5-Turbo is a 8x7B model, basically the same size as Mixtral. (this has changed and most likely not true, see update section).

One thing I also want to note is that the paper I mentioned was a newer paper compared to one about a month or two ago that described a similar technique to this? (I forgot the name of that earlier paper). They did something similar and found the embedding size of other smaller OpenAI models, but they did not give GPT-3.5-Turbo's embedding size as per request of OpenAI. (heavy speculation ahead) Those paper authors not giving that information might be due to a model with the same specifications as GPT-3.5-Turbo already existing, and that model very much might actually be Mixtral!

I do want to point out that we do not know the size of the model without these estimates, but if we go completely off the paper without any of this extra reasoning and speculation (not counting a dense 7B as it is not big enough to explain a lot of the performance with GPT-3.5-Turbo), GPT-3.5-Turbo would be a Nx7±\~2B MoE (This shorthand I am using looks awful, but basically N number of experts (unknown, can't be related to embedding size) with 7B for each expert, give or take around 2B parameters for each expert (embedding size of 4096 ± 512). (new estimates found below)

**UPDATE**: The ratio between parameter size and embedding size of different models is sort of exponential as parameter size scales up. If we take GPT-2 and GPT-3's embedding and parameter count and go off of that, we can get a better idea of the parameter size for each expert in GPT-3.5-Turbo. (BIG thanks to u/[NighthawkT42](https://www.reddit.com/user/NighthawkT42/) for pointing me in a better direction for this prediction and giving a better size prediction of GPT-3.5-Turbo based more in calculation!)

GPT-2-Small : 124M parameters, 768 embedding size, ratio of 161,458

GPT-2-XL : 1.5B parameters, 1,600 embedding size, ratio of 937,500

GPT-3-175B : 175B parameters, 12,288 embedding size, ratio of 14,241,536

By fitting the curve to these data points, we get these predicted values for parameter size:

https://preview.redd.it/c5akcngvalsc1.png?width=800&format=png&auto=webp&s=4b22892c709300a5f35f71ac94d1c002aa1449b1

3584 embedding = \~8.3B parameters
4096 embedding = \~11.6B parameters
4608 embedding = \~15.7B parameters

By using this new evidence, a 8x7 model size, while still possible when compared to other model embedding sizes compared to their parameters, is now not really likely when estimating from these values from older OpenAI models. While of course they are still estimates, if we take the average of the 4608 embedding model and the 3584 embedding model, we get a size of \~12B parameters for each expert.

If we follow along with 8 experts still, it is possible that GPT-3.5-Turbo is a 8x12B MoE. This makes a lot more sense in why it is stronger at more diverse languages compared to Mixtral, having a bigger impact than training data as my old prediction above states.

If we remove all of this speculation and just take the calculated values and the values from the paper (and assume not a dense model), then **GPT-3.5-Turbo is a Nx12B±\~4B** (N number of experts, 12B parameters each with a possible deviation of \~4B parameters).

Open Reddit thread
r/just4ochat 14 upvotes 5 comments March 9, 2026
Should we add GPT-3.5 Turbo to the platform?

I was just clicking around and I saw GPT-3.5 is still live on the API!

For those who remember those days, should we add it into just4o?

It may not work with some of our more advanced features like image generation tools or BLS/EIA statistics, but it would be a fun add for folks who want a taste of OG AI.

Link to the GPT-3.5 turbo dev page for more info: https://developers.openai.com/api/docs/models/gpt-3.5-turbo

We’re curious about your thoughts! 💚

Open Reddit thread
View more discussions →

More models from OpenAI

Continue browsing adjacent models from the same provider.

← All AI Models