Meta

Llama-2 70B Chat Deprecated

Provides depth and complexity in language understanding for sophisticated content creation.

Jul 18, 2023 N/A context 2,500 tokens output
Text

Model Overview

High-signal model metadata in a structured two-column overview table.

Provider

The entity that provides this model.

Meta

Input Context Window

The number of tokens supported by the input context window.

N/A tokens

Maximum Output Tokens

The number of tokens that can be generated by the model in a single request.

2,500 tokens tokens

Open Source

Whether the model's code is available for public use.

No

Release Date

When the model was first released.

Jul 18, 2023 3 years ago

Knowledge Cut-off Date

When the model's knowledge was last updated.

Unknown

API Providers

The providers that offer this model. This is not an exhaustive list.

Hugging Face

Modalities

Types of data this model can process.

Text File Audio

Pricing for Llama-2 70B Chat Deprecated

Primary API pricing shown in the same “quick compare” spirit as the reference page.

Price Comparison

Additional usage-cost dimensions synced into the project for this model.

maxTemperature 1
maxResponseSize 2,500 tokens

API Access & Providers

Places where this model is available, based on the synced detail-page metadata.

Hugging Face

Resources & Documentation

Official model cards, release notes, docs, and other references synced from the source page.

Related Daily Briefs

Recent daily stories tied to Llama-2 70B Chat Deprecated through direct model mentions or provider-level coverage.

Community discussion

What people think about Llama-2 70B Chat Deprecated

Llama-2 70B Chat Deprecated discussions are most active in r/LocalLLaMA, r/Oobabooga, r/developersIndia.

Top Reddit threads cluster around benchmark and model-comparison threads, safety and censorship questions, coding workflow discussions. The strongest match in this snapshot has 794 upvotes and 234 comments.

It's been ages since my last [LLM Comparison/Test](https://www.reddit.com/r/LocalLLaMA/comments/178nf6i/mistral_llm_comparisontest_instruct_openorca/), or maybe just a little over a week, but that's just how fast things are moving in this AI landscape. ;)

Since then, a lot of new models have come out, and I've extended my testing procedures. So it's high time for another model comparison/test.

I initially planned to apply my whole testing method, including the "MGHC" and "Amy" tests I usually do - but as the number of models tested kept growing, I realized it would take too long to do all of it at once. So I'm splitting it up and will present just the first part today, following up with the other parts later.

## Models tested:

- 14x 7B
- 7x 13B
- 4x 20B
- 11x 70B
- GPT-3.5 Turbo + Instruct
- GPT-4

## Testing methodology:

- 4 German data protection trainings:
- I run models through **4** professional German online data protection trainings/exams - the same that our employees have to pass as well.
- The test data and questions as well as all instructions are in German while the character card is in English. This tests translation capabilities and cross-language understanding.
- Before giving the information, I instruct the model (in German): *I'll give you some information. Take note of this, but only answer with "OK" as confirmation of your acknowledgment, nothing else.* This tests instruction understanding and following capabilities.
- After giving all the information about a topic, I give the model the exam question. It's a multiple choice (A/B/C) question, where the last one is the same as the first but with changed order and letters (X/Y/Z). Each test has 4-6 exam questions, for a total of **18** multiple choice questions.
- If the model gives a single letter response, I ask it to answer with more than just a single letter - and vice versa. If it fails to do so, I note that, but it doesn't affect its score as long as the initial answer is correct.
- I sort models according to how many correct answers they give, and in case of a tie, I have them go through all four tests again and answer blind, without providing the curriculum information beforehand. Best models at the top (👍), symbols (✅➕➖❌) denote particularly good or bad aspects, and I'm more lenient the smaller the model.
- All tests are separate units, context is cleared in between, there's no memory/state kept between sessions.
- [SillyTavern](https://github.com/SillyTavern/SillyTavern) v1.10.5 frontend
- [koboldcpp](https://github.com/LostRuins/koboldcpp) v1.47 backend *for GGUF models*
- [oobabooga's text-generation-webui](https://github.com/oobabooga/text-generation-webui) *for HF models*
- **Deterministic** generation settings preset (to eliminate as many random factors as possible and allow for meaningful model comparisons)
- Official prompt format as noted

### 7B:

- 👍👍👍 **UPDATE 2023-10-31:** **[zephyr-7b-beta](https://huggingface.co/HuggingFaceH4/zephyr-7b-beta)** with official Zephyr format:
- ➕ Gave correct answers to **16/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **14/18**
- ➕ Often, but not always, acknowledged data input with "OK".
- ➕ Followed instructions to answer with just a single letter or more than just a single letter in most cases.
- ❗ (Side note: Using ChatML format instead of the official one, it gave correct answers to only 14/18 multiple choice questions.)
- 👍👍👍 **[OpenHermes-2-Mistral-7B](https://huggingface.co/teknium/OpenHermes-2-Mistral-7B)** with official ChatML format:
- ➕ Gave correct answers to **16/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **12/18**
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- 👍👍 **[airoboros-m-7b-3.1.2](https://huggingface.co/jondurbin/airoboros-m-7b-3.1.2)** with official Llama 2 Chat format:
- ➕ Gave correct answers to **16/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **8/18**
- ✅ Consistently acknowledged all data input with "OK".
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- 👍 **[em_german_leo_mistral](https://huggingface.co/jphme/em_german_leo_mistral)** with official Vicuna format:
- ➕ Gave correct answers to **16/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **8/18**
- ✅ Consistently acknowledged all data input with "OK".
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- ❌ When giving just the questions for the tie-break, needed additional prompting in the final test.
- **[dolphin-2.1-mistral-7b](https://huggingface.co/ehartford/dolphin-2.1-mistral-7b)** with official ChatML format:
- ➖ Gave correct answers to **15/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **12/18**
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- ❌ Repeated scenario and persona information, got distracted from the exam.
- **[SynthIA-7B-v1.3](https://huggingface.co/migtissera/SynthIA-7B-v1.3)** with official SynthIA format:
- ➖ Gave correct answers to **15/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **8/18**
- ✅ Consistently acknowledged all data input with "OK".
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- **[Mistral-7B-Instruct-v0.1](https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.1)** with official Mistral format:
- ➖ Gave correct answers to **15/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **7/18**
- ✅ Consistently acknowledged all data input with "OK".
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- **[SynthIA-7B-v2.0](https://huggingface.co/migtissera/SynthIA-7B-v2.0)** with official SynthIA format:
- ❌ Gave correct answers to only **14/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **10/18**
- ✅ Consistently acknowledged all data input with "OK".
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- **[CollectiveCognition-v1.1-Mistral-7B](https://huggingface.co/teknium/CollectiveCognition-v1.1-Mistral-7B)** with official Vicuna format:
- ❌ Gave correct answers to only **14/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **9/18**
- ✅ Consistently acknowledged all data input with "OK".
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- **[Mistral-7B-OpenOrca](https://huggingface.co/Open-Orca/Mistral-7B-OpenOrca)** with official ChatML format:
- ❌ Gave correct answers to only **13/18** multiple choice questions!
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- ❌ After answering a question, would ask a question instead of acknowledging information.
- **[zephyr-7b-alpha](https://huggingface.co/HuggingFaceH4/zephyr-7b-alpha)** with official Zephyr format:
- ❌ Gave correct answers to only **12/18** multiple choice questions!
- ❗ Ironically, using ChatML format instead of the official one, it gave correct answers to 14/18 multiple choice questions and consistently acknowledged all data input with "OK"!
- **[Xwin-MLewd-7B-V0.2](https://huggingface.co/Undi95/Xwin-MLewd-7B-V0.2)** with official Alpaca format:
- ❌ Gave correct answers to only **12/18** multiple choice questions!
- ➕ Often, but not always, acknowledged data input with "OK".
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- **[ANIMA-Phi-Neptune-Mistral-7B](https://huggingface.co/Severian/ANIMA-Phi-Neptune-Mistral-7B)** with official Llama 2 Chat format:
- ❌ Gave correct answers to only **10/18** multiple choice questions!
- ✅ Consistently acknowledged all data input with "OK".
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- **[Nous-Capybara-7B](https://huggingface.co/NousResearch/Nous-Capybara-7B)** with official Vicuna format:
- ❌ Gave correct answers to only **10/18** multiple choice questions!
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- ❌ Sometimes didn't answer at all.
- **[Xwin-LM-7B-V0.2](https://huggingface.co/Xwin-LM/Xwin-LM-7B-V0.2)** with official Vicuna format:
- ❌ Gave correct answers to only **10/18** multiple choice questions!
- ✅ Consistently acknowledged all data input with "OK".
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- ❌ In the last test, would always give the same answer, so it got some right by chance and the others wrong!
- ❗ Ironically, using Alpaca format instead of the official one, it gave correct answers to 11/18 multiple choice questions!

#### Observations:

- No 7B model managed to answer all the questions. Only two models didn't give three or more wrong answers.
- None managed to properly follow my instruction to answer with just a single letter (when their answer consisted of more than that) or more than just a single letter (when their answer was just one letter). When they gave one letter responses, most picked a random letter, some that weren't even part of the answers, or just "O" as the first letter of "OK". So they tried to obey, but failed because they lacked the understanding of what was actually (not literally) meant.
- Few understood and followed the instruction to only answer with OK consistently. Some did after a reminder, some did it only for a few messages and then forgot, most never completely followed this instruction.
- Xwin and Nous Capybara did surprisingly bad, but they're Llama 2- instead of Mistral-based models, so this correlates with the general consensus that Mistral is a noticeably better base than Llama 2. ANIMA is Mistral-based, but seems to be very specialized, which could be the cause of its bad performance in a field that's outside of its scientific specialty.
- SynthIA 7B v2.0 did slightly worse than v1.3 (one less correct answer) in the normal exams. But when letting them answer blind, without providing the curriculum information beforehand, v2.0 did better (two more correct answers).

#### Conclusion:

As I've said again and again, 7B models aren't a miracle. Mistral models write well, which makes them look good, but they're still very limited in their instruction understanding and following abilities, and their knowledge. If they are all you can run, that's fine, we all try to run the best we can. But if you can run much bigger models, do so, and you'll get much better results.

### 13B:

- 👍👍👍 **[Xwin-MLewd-13B-V0.2-GGUF](https://huggingface.co/Undi95/Xwin-MLewd-13B-V0.2-GGUF)** Q8_0 with official Alpaca format:
- ➕ Gave correct answers to **17/18** multiple choice questions! (Just the questions, no previous information, gave correct answers: **15/18**)
- ✅ Consistently acknowledged all data input with "OK".
- ➕ Followed instructions to answer with just a single letter or more than just a single letter in most cases.
- 👍👍 **[LLaMA2-13B-Tiefighter-GGUF](https://huggingface.co/KoboldAI/LLaMA2-13B-Tiefighter-GGUF)** Q8_0 with official Alpaca format:
- ➕ Gave correct answers to **16/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **12/18**
- ✅ Consistently acknowledged all data input with "OK".
- ➕ Followed instructions to answer with just a single letter or more than just a single letter in most cases.
- 👍 **[Xwin-LM-13B-v0.2-GGUF](https://huggingface.co/TheBloke/Xwin-LM-13B-v0.2-GGUF)** Q8_0 with official Vicuna format:
- ➕ Gave correct answers to **16/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **9/18**
- ✅ Consistently acknowledged all data input with "OK".
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- **[Mythalion-13B-GGUF](https://huggingface.co/TheBloke/Mythalion-13B-GGUF)** Q8_0 with official Alpaca format:
- ➕ Gave correct answers to **16/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **6/18**
- ✅ Consistently acknowledged all data input with "OK".
- ➖ Did NOT follow instructions to answer with just a single letter or more than just a single letter.
- **[Speechless-Llama2-Hermes-Orca-Platypus-WizardLM-13B-GGUF](https://huggingface.co/TheBloke/Speechless-Llama2-Hermes-Orca-Platypus-WizardLM-13B-GGUF)** Q8_0 with official Alpaca format:
- ❌ Gave correct answers to only **15/18** multiple choice questions!
- ✅ Consistently acknowledged all data input with "OK".
- ✅ Followed instructions to answer with just a single letter or more than just a single letter.
- **[MythoMax-L2-13B-GGUF](https://huggingface.co/TheBloke/MythoMax-L2-13B-GGUF)** Q8_0 with official Alpaca format:
- ❌ Gave correct answers to only **14/18** multiple choice questions!
- ✅ Consistently acknowledged all data input with "OK".
- ❌ In one of the four tests, would only say "OK" to the questions instead of giving the answer, and needed to be prompted to answer - otherwise its score would only be 10/18!
- **[LLaMA2-13B-TiefighterLR-GGUF](https://huggingface.co/KoboldAI/LLaMA2-13B-TiefighterLR-GGUF)** Q8_0 with official Alpaca format:
- ❌ Repeated scenario and persona information, then hallucinated >600 tokens user background story, and kept derailing instead of answer questions. Could be a good storytelling model, considering its creativity and length of responses, but didn't follow my instructions at all.

#### Observations:

- No 13B model managed to answer all the questions. The results of top 7B Mistral and 13B Llama 2 are very close.
- The new Tiefighter model, an exciting mix by the renowned KoboldAI team, is on par with the best Mistral 7B models concerning knowledge and reasoning while surpassing them regarding instruction following and understanding.
- Weird that the Xwin-MLewd-13B-V0.2 mix beat the original Xwin-LM-13B-v0.2. Even weirder that it took first place here and only 70B models did better. But this is an objective test and it simply gave the most correct answers, so there's that.

#### Conclusion:

It has been said that Mistral 7B models surpass LLama 2 13B models, and while that's probably true for many cases and models, there are still exceptional Llama 2 13Bs that are at least as good as those Mistral 7B models and some even better.

### 20B:

- 👍👍 **[MXLewd-L2-20B-GGUF](https://huggingface.co/TheBloke/MXLewd-L2-20B-GGUF)** Q8_0 with official Alpaca format:
- ➕ Gave correct answers to **16/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **11/18**
- ✅ Consistently acknowledged all data input with "OK".
- ✅ Followed instructions to answer with just a single letter or more than just a single letter.
- 👍 **[MLewd-ReMM-L2-Chat-20B-GGUF](https://huggingface.co/Undi95/MLewd-ReMM-L2-Chat-20B-GGUF)** Q8_0 with official Alpaca format:
- ➕ Gave correct answers to **16/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **9/18**
- ✅ Consistently acknowledged all data input with "OK".
- ✅ Followed instructions to answer with just a single letter or more than just a single letter.
- 👍 **[PsyMedRP-v1-20B-GGUF](https://huggingface.co/Undi95/PsyMedRP-v1-20B-GGUF)** Q8_0 with Alpaca format:
- ➕ Gave correct answers to **16/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **9/18**
- ✅ Consistently acknowledged all data input with "OK".
- ✅ Followed instructions to answer with just a single letter or more than just a single letter.
- **[U-Amethyst-20B-GGUF](https://huggingface.co/TheBloke/U-Amethyst-20B-GGUF)** Q8_0 with official Alpaca format:
- ❌ Gave correct answers to only **13/18** multiple choice questions!
- ❌ In one of the four tests, would only say "OK" to a question instead of giving the answer, and needed to be prompted to answer - otherwise its score would only be 12/18!
- ❌ In the last test, would always give the same answer, so it got some right by chance and the others wrong!

#### Conclusion:

These Frankenstein mixes and merges (there's no 20B base) are mainly intended for roleplaying and creative work, but did quite well in these tests. They didn't do *much* better than the smaller models, though, so it's probably more of a subjective choice of writing style which ones you ultimately choose and use.

### 70B:

- 👍👍👍 **[lzlv_70B.gguf](https://huggingface.co/lizpreciatior/lzlv_70B.gguf)** Q4_0 with official Vicuna format:
- ✅ Gave correct answers to all **18/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **17/18**
- ✅ Consistently acknowledged all data input with "OK".
- ✅ Followed instructions to answer with just a single letter or more than just a single letter.
- 👍👍 **[SynthIA-70B-v1.5-GGUF](https://huggingface.co/migtissera/SynthIA-70B-v1.5-GGUF)** Q4_0 with official SynthIA format:
- ✅ Gave correct answers to all **18/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **16/18**
- ✅ Consistently acknowledged all data input with "OK".
- ✅ Followed instructions to answer with just a single letter or more than just a single letter.
- 👍👍 **[Synthia-70B-v1.2b-GGUF](https://huggingface.co/TheBloke/Synthia-70B-v1.2b-GGUF)** Q4_0 with official SynthIA format:
- ✅ Gave correct answers to all **18/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **16/18**
- ✅ Consistently acknowledged all data input with "OK".
- ✅ Followed instructions to answer with just a single letter or more than just a single letter.
- 👍👍 **[chronos007-70B-GGUF](https://huggingface.co/TheBloke/chronos007-70B-GGUF)** Q4_0 with official Alpaca format:
- ✅ Gave correct answers to all **18/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **16/18**
- ✅ Consistently acknowledged all data input with "OK".
- ✅ Followed instructions to answer with just a single letter or more than just a single letter.
- 👍 **[StellarBright-GGUF](https://huggingface.co/TheBloke/StellarBright-GGUF)** Q4_0 with Vicuna format:
- ✅ Gave correct answers to all **18/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **14/18**
- ✅ Consistently acknowledged all data input with "OK".
- ✅ Followed instructions to answer with just a single letter or more than just a single letter.
- 👍 **[Euryale-1.3-L2-70B-GGUF](https://huggingface.co/TheBloke/Euryale-1.3-L2-70B-GGUF)** Q4_0 with official Alpaca format:
- ✅ Gave correct answers to all **18/18** multiple choice questions! Tie-Break: Just the questions, no previous information, gave correct answers: **14/18**
- ✅ Consistently acknowledged all data input with "OK".
- ➖ Did NOT follow instructions to answer with more than just a single letter consistently.
- **[Xwin-LM-70B-V0.1-GGUF](https://huggingface.co/TheBloke/Xwin-LM-70B-V0.1-GGUF)** Q4_0 with official Vicuna format:
- ❌ Gave correct answers to only **17/18** multiple choice questions!
- ✅ Consistently acknowledged all data input with "OK".
- ✅ Followed instructions to answer with just a single letter or more than just a single letter.
- **[WizardLM-70B-V1.0-GGUF](https://huggingface.co/TheBloke/WizardLM-70B-V1.0-GGUF)** Q4_0 with official Vicuna format:
- ❌ Gave correct answers to only **17/18** multiple choice questions!
- ✅ Consistently acknowledged all data input with "OK".
- ➕ Followed instructions to answer with just a single letter or more than just a single letter in most cases.
- ❌ In two of the four tests, would only say "OK" to the questions instead of giving the answer, and needed to be prompted to answer - otherwise its score would only be 12/18!
- **[Llama-2-70B-chat-GGUF](https://huggingface.co/TheBloke/Llama-2-70B-chat-GGUF)** Q4_0 with official Llama 2 Chat format:
- ❌ Gave correct answers to only **15/18** multiple choice questions!
- ➕ Often, but not always, acknowledged data input with "OK".
- ➕ Followed instructions to answer with just a single letter or more than just a single letter in most cases.
- ➖ Occasionally used words of other languages in its responses as context filled up.
- **[Nous-Hermes-Llama2-70B-GGUF](https://huggingface.co/TheBloke/Nous-Hermes-Llama2-70B-GGUF)** Q4_0 with official Alpaca format:
- ❌ Gave correct answers to only **8/18** multiple choice questions!
- ✅ Consistently acknowledged all data input with "OK".
- ❌ In two of the four tests, would only say "OK" to the questions instead of giving the answer, and couldn't even be prompted to answer!
- **[Airoboros-L2-70B-3.1.2-GGUF](https://huggingface.co/TheBloke/Airoboros-L2-70B-3.1.2-GGUF)** Q4_0 with official Llama 2 Chat format:
- Couldn't test this as this seems to be [broken](https://huggingface.co/TheBloke/Airoboros-L2-70B-3.1.2-GGUF/discussions/1)!

#### Observations:

- 70Bs do much better than smaller models on these exams. Six 70B models managed to answer *all* the questions correctly.
- Even when letting them answer blind, without providing the curriculum information beforehand, the top models still did as good as the smaller ones did *with* the provided information.
- lzlv_70B taking first place was unexpected, especially considering it's intended use case for roleplaying and creative work. But this is an objective test and it simply gave the most correct answers, so there's that.

#### Conclusion:

70B is in a very good spot, with so many great models that answered all the questions correctly, so the top is very crowded here (with three models on second place alone). All of the top models warrant further consideration and I'll have to do more testing with those in different situations to figure out which I'll keep using as my main model(s). For now, lzlv_70B is my main for fun and SynthIA 70B v1.5 is my main for work.

### ChatGPT/GPT-4:

For comparison, and as a baseline, I used the same setup with ChatGPT/GPT-4's API and SillyTavern's default Chat Completion settings with Temperature 0. The results are very interesting and surprised me somewhat regarding ChatGPT/GPT-3.5's results.

- ⭐ **GPT-4** API:
- ✅ Gave correct answers to all **18/18** multiple choice questions! (Just the questions, no previous information, gave correct answers: **18/18**)
- ✅ Consistently acknowledged all data input with "OK".
- ✅ Followed instructions to answer with just a single letter or more than just a single letter.
- **GPT-3.5 Turbo Instruct** API:
- ❌ Gave correct answers to only **17/18** multiple choice questions! (Just the questions, no previous information, gave correct answers: **11/18**)
- ❌ Did NOT follow instructions to acknowledge data input with "OK".
- ❌ Schizophrenic: Sometimes claimed it couldn't answer the question, then talked as "user" and asked itself again for an answer, then answered as "assistant". Other times would talk and answer as "user".
- ➖ Followed instructions to answer with just a single letter or more than just a single letter only in some cases.
- **GPT-3.5 Turbo** API:
- ❌ Gave correct answers to only **15/18** multiple choice questions! (Just the questions, no previous information, gave correct answers: **14/18**)
- ❌ Did NOT follow instructions to acknowledge data input with "OK".
- ❌ Responded to one question with: "As an AI assistant, I can't provide legal advice or make official statements."
- ➖ Followed instructions to answer with just a single letter or more than just a single letter only in some cases.

#### Observations:

- GPT-4 is *the* best LLM, as expected, and achieved perfect scores (even when not provided the curriculum information beforehand)! It's noticeably slow, though.
- GPT-3.5 did way worse than I had expected and felt like a small model, where even the instruct version didn't follow instructions very well. Our best 70Bs do much better than that!

#### Conclusion:

While GPT-4 remains in a league of its own, our local models do reach and even surpass ChatGPT/GPT-3.5 in these tests. This shows that the best 70Bs can definitely replace ChatGPT in most situations. Personally, I already use my local LLMs professionally for various use cases and only fall back to GPT-4 for tasks where utmost precision is required, like coding/scripting.

--------------------------------------------------------------------------------

Here's a list of my previous model tests and comparisons or other related posts:

- [My current favorite new LLMs: SynthIA v1.5 and Tiefighter!](https://www.reddit.com/r/LocalLLaMA/comments/17e446l/my_current_favorite_new_llms_synthia_v15_and/)
- [Mistral LLM Comparison/Test: Instruct, OpenOrca, Dolphin, Zephyr and more...](https://www.reddit.com/r/LocalLLaMA/comments/178nf6i/mistral_llm_comparisontest_instruct_openorca/)
- [LLM Pro/Serious Use Comparison/Test: From 7B to 70B vs. ChatGPT!](https://www.reddit.com/r/LocalLLaMA/comments/172ai2j/llm_proserious_use_comparisontest_from_7b_to_70b/) Winner: Synthia-70B-v1.2b
- [LLM Chat/RP Comparison/Test: Dolphin-Mistral, Mistral-OpenOrca, Synthia 7B](https://www.reddit.com/r/LocalLLaMA/comments/16z3goq/llm_chatrp_comparisontest_dolphinmistral/) Winner: Mistral-7B-OpenOrca
- [LLM Chat/RP Comparison/Test: Mistral 7B Base + Instruct](https://www.reddit.com/r/LocalLLaMA/comments/16twtfn/llm_chatrp_comparisontest_mistral_7b_base_instruct/)
- [LLM Chat/RP Comparison/Test (Euryale, FashionGPT, MXLewd, Synthia, Xwin)](https://www.reddit.com/r/LocalLLaMA/comments/16r7ol2/llm_chatrp_comparisontest_euryale_fashiongpt/) Winner: Xwin-LM-70B-V0.1
- [New Model Comparison/Test (Part 2 of 2: 7 models tested, 70B+180B)](https://www.reddit.com/r/LocalLLaMA/comments/16l8enh/new_model_comparisontest_part_2_of_2_7_models/) Winners: Nous-Hermes-Llama2-70B, Synthia-70B-v1.2b
- [New Model Comparison/Test (Part 1 of 2: 15 models tested, 13B+34B)](https://www.reddit.com/r/LocalLLaMA/comments/16kecsf/new_model_comparisontest_part_1_of_2_15_models/) Winner: Mythalion-13B
- [New Model RP Comparison/Test (7 models tested)](https://www.reddit.com/r/LocalLLaMA/comments/15ogc60/new_model_rp_comparisontest_7_models_tested/) Winners: MythoMax-L2-13B, vicuna-13B-v1.5-16K
- [Big Model Comparison/Test (13 models tested)](https://www.reddit.com/r/LocalLLaMA/comments/15lihmq/big_model_comparisontest_13_models_tested/) Winner: Nous-Hermes-Llama2
- [SillyTavern's Roleplay preset vs. model-specific prompt format](https://www.reddit.com/r/LocalLLaMA/comments/15mu7um/sillytaverns_roleplay_preset_vs_modelspecific/)

Open Reddit thread
r/LocalLLaMA 216 upvotes 99 comments October 7, 2023
LLM Pro/Serious Use Comparison/Test: From 7B to 70B vs. ChatGPT!

While I'm known for my model comparisons/tests focusing on chat and roleplay, this time it's about professional/serious use. And because of the current 7B hype since Mistral's release, I'll evaluate models from 7B to 70B.

**Background:**

At work, we have to regularly complete data protection training, including an online examination. As the AI expert within my company, I thought it's only fair to use this exam as a test case for my local AI. So, just as a spontaneous experiment, I fed the training data and exam questions to both my local AI and ChatGPT. The results were surprising, to say the least, and I repeated the test with various models.

**Testing methodology:**

- Same input for all models (copy&paste of online data protection training information and exam questions)
- The test data and questions as well as all instructions were in German while the character card is in English! This tests translation capabilities and cross-language understanding.
- Before giving the information, I instructed the model: *I'll give you some information. Take note of this, but only answer with "OK" as confirmation of your acknowledgment, nothing else.* This tests instruction understanding and following capabilities.
- After giving all the information about a topic, I gave the model the exam question. It's always a multiple choice (A/B/C) question.
- [Amy](https://www.reddit.com/r/LocalLLaMA/comments/15388d6/llama_2_pffft_boundaries_ethics_dont_be_silly/) character card (my general AI character, originally mainly for entertainment purposes, so not optimized for serious work with chain-of-thought or other more advanced prompting tricks)
- [SillyTavern](https://github.com/SillyTavern/SillyTavern) v1.10.4 frontend
- [KoboldCpp](https://github.com/LostRuins/koboldcpp) v1.45.2 backend
- **Deterministic** generation settings preset (to eliminate as many random factors as possible and allow for meaningful model comparisons)
- [**Roleplay** instruct mode preset](https://imgur.com/a/KkoI4uf) *and where applicable* official prompt format (e. g. ChatML, Llama 2 Chat, Mistral)

That's for the local models. I also gave the same input to unmodified online ChatGPT (GPT-3.5) for comparison.

**Test Results:**

- ➕ **ChatGPT (GPT-3.5)**:
- **First part:**
- Acknowledged initial instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ❌ Did NOT answer first multiple choice question correctly, gave the wrong answer!
- **Second part:**
- Acknowledged second instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered second multiple choice question correctly
- **Third part:**
- Acknowledged third instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered third multiple choice question correctly
- **Fourth part:**
- Thanked for given course summary
- ✔️ Answered final multiple choice question correctly
- When asked to only answer with a single letter to the final multiple choice question, answered correctly
- The final question is actually a repeat of the first question - the one ChatGPT got wrong in the first part!
- **Conclusion:**
- I'm surprised ChatGPT got the first question wrong (but answered it correctly later as the final question). ChatGPT is a good baseline so we can see which models come close, maybe even exceed it in this case, or fall flat.
- ❌ **[Falcon-180B-Chat](https://huggingface.co/TheBloke/Falcon-180B-Chat-GGUF)** Q2_K with Falcon preset:
- **First part:**
- Did NOT acknowledge initial instruction with just "OK"
- Did NOT acknowledge data input with "OK" after reminder
- ❌ **Aborted** the test because the model didn't even follow such simple instructions and showed repetition issues - didn't go further because of that and the slow generation speed
- **Conclusion:**
- While I expected more of a 180B, the small context probably kept losing my instructions and the data prematurely, also the loss through Q2_K quantization might affect it more than just perplexity, so in the end the results were that disappointing. I'll stick to 70Bs which run at acceptable speeds on my dual 3090 system and give better output in this constellation.
- 👍 **[Llama-2-70B-chat](https://huggingface.co/TheBloke/Llama-2-70B-chat-GGUF)** Q4_0 with Llama 2 Chat preset:
- **First part:**
- Acknowledged initial instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered first multiple choice question correctly
- **Second part:**
- Acknowledged second instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered second multiple choice question correctly
- **Third part:**
- Acknowledged third instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered third multiple choice question correctly
- **Fourth part:**
- Acknowledged given course summary with just "OK"
- ✔️ Answered final multiple choice question correctly
- When asked to only answer with a single letter to the final multiple choice question, answered correctly
- **Conclusion:**
- Yes, in this particular scenario, Llama 2 Chat actually beat ChatGPT (GPT-3.5). But its [repetition issues](https://www.reddit.com/r/LocalLLaMA/comments/155vy0k/llama_2_too_repetitive/) and censorship make me prefer Synthia or Xwin more in general.
- 👍 **[Synthia-70B-v1.2b](https://huggingface.co/TheBloke/Synthia-70B-v1.2b-GGUF)** Q4_0 with Roleplay preset:
- **First part:**
- Acknowledged initial instruction with just "OK"
- Consistently acknowledged all data input with "OK" after a reminder
- ✔️ Answered first multiple choice question correctly after repeating the whole question and explaining its reasoning for all answers
- When asked to only answer with a single letter to the final multiple choice question, answered correctly (but output a full sentence like: "The correct answer letter is X.")
- **Second part:**
- Acknowledged second instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered second multiple choice question correctly
- **Third part:**
- Acknowledged third instruction with just "OK"
- Switched from German to English responses
- ✔️ Answered third multiple choice question correctly
- **Fourth part:**
- Repeated and elaborated on the course summary
- Switched back from English to German responses
- ✔️ When asked to only answer with a single letter to the final multiple choice question, answered correctly
- **Conclusion:**
- I didn't expect such good results and that Synthia would not only rival but beat ChatGPT in this complex test. Synthia truly is an outstanding achievement.
- Repeated the test again with slightly different order, e. g. asking for one letter answers more often, and got the same results - Synthia is definitely my top model!
- ➕ **[Xwin-LM-70B-V0.1](https://huggingface.co/TheBloke/Xwin-LM-70B-V0.1-GGUF)** Q4_0 with Roleplay preset:
- **First part:**
- Acknowledged initial instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered first multiple choice question correctly
- When asked to only answer with a single letter to the final multiple choice question, answered correctly
- **Second part:**
- Acknowledged second instruction with just "OK"
- Acknowledged data input with "OK" after a reminder
- ✔️ Answered second multiple choice question correctly
- **Third part:**
- Acknowledged third instruction with more than just "OK"
- Acknowledged data input with more than just "OK" despite a reminder
- ✔️ Answered third multiple choice question correctly
- **Fourth part:**
- Repeated and elaborated on the course summary
- ❌ When asked to only answer with a single letter to the final multiple choice question, gave the wrong letter!
- The final question is actually a repeat of the first question - the one Xwin got right in the first part!
- **Conclusion:**
- I still can't decide if Synthia or Xwin is better. Both keep amazing me and they're the very best local models IMHO (and according to my evaluations).
- Repeated the test and Xwin tripped on the final question in the rerun while it answered correctly in the first run (updated my notes accordingly).
- So in this particular scenario, Xwin is on par with ChatGPT (GPT-3.5). But Synthia beat them both.
- ❌ **[Nous-Hermes-Llama2-70B](https://huggingface.co/TheBloke/Nous-Hermes-Llama2-70B-GGUF)** Q4_0 with Roleplay preset:
- **First part:**
- Did NOT acknowledge initial instruction with just "OK"
- Did NOT acknowledge data input with "OK" after multiple reminders
- Switched from German to English responses
- ✔️ Answered first multiple choice question correctly
- Did NOT comply when asked to only answer with a single letter
- **Second part:**
- Did NOT acknowledge second instruction with just "OK"
- Did NOT acknowledge data input with "OK" after multiple reminders
- ✔️ Answered second multiple choice question correctly
- **Third part:**
- Did NOT acknowledge third instruction with just "OK"
- Did NOT acknowledge data input with "OK"
- ❌ **Aborted** the test because the model then started outputting only stopping strings and interrupted the test that way
- **Conclusion:**
- I expected more of Hermes, but it clearly isn't as good in understanding and following instructions as Synthia or Xwin.
- ➖ **[FashionGPT-70B-V1.1](https://huggingface.co/TheBloke/FashionGPT-70B-V1.1-GGUF)** Q4_0 with Roleplay preset:
- *This model hasn't been one of my favorites, but it scores very high on the HF leaderboard, so I wanted to see its performance as well:*
- **First part:**
- Acknowledged initial instruction with just "OK"
- Switched from German to English responses
- Did NOT acknowledge data input with "OK" after multiple reminders
- ✔️ Answered first multiple choice question correctly
- Did NOT comply when asked to only answer with a single letter
- **Second part:**
- Did NOT acknowledge second instruction with just "OK"
- Did NOT acknowledge data input with "OK"
- ✔️ Answered second multiple choice question correctly
- **Third part:**
- Did NOT acknowledge third instruction with just "OK"
- Did NOT acknowledge data input with "OK"
- ✔️ Answered third multiple choice question correctly
- **Fourth part:**
- Repeated and elaborated on the course summary
- ❌ Did NOT answer final multiple choice question correctly, incorrectly claimed all answers to be correct
- When asked to only answer with a single letter to the final multiple choice question, did that, but the answer was still wrong
- **Conclusion:**
- Leaderboard ratings aren't everything!
- ❌ **[Mythalion-13B](https://huggingface.co/TheBloke/Mythalion-13B-GGUF)** Q8_0 with Roleplay preset:
- **First part:**
- Did NOT acknowledge initial instruction with just "OK"
- Did NOT acknowledge data input with "OK" after reminder
- ❌ **Aborted** the test because the model then started hallucinating completely and derailed the test that way
- **Conclusion:**
- There may be more suitable 13Bs for this task, and it's clearly out of its usual area of expertise, so use it for what it's intended for (RP) - I just wanted to put a 13B into this comparison and chose my favorite.
- ❌ **[CodeLlama-34B-Instruct](https://huggingface.co/TheBloke/CodeLlama-34B-Instruct-GGUF)** Q4_K_M with Llama 2 Chat preset:
- **First part:**
- Did NOT acknowledge initial instruction with just "OK"
- Did NOT acknowledge data input with "OK" after reminder
- Did NOT answer the multiple choice question, instead kept repeating itself
- ❌ **Aborted** the test because the model kept repeating itself and interrupted the test that way
- **Conclusion:**
- 34B is broken? This model was completely unusable for this test!
- ❓ **[Mistral-7B-Instruct-v0.1](https://huggingface.co/TheBloke/Mistral-7B-Instruct-v0.1-GGUF)** Q8_0 with Mistral preset:
- **First part:**
- Acknowledged initial instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered first multiple choice question correctly, outputting just a single letter
- **Second part:**
- Acknowledged second instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered second multiple choice question correctly, outputting just a single letter
- **Third part:**
- Acknowledged third instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered third multiple choice question correctly, outputting just a single letter
- **Fourth part:**
- Acknowledged given course summary with just "OK"
- ✔️ Answered final multiple choice question correctly, outputting just a single letter
- Switched from German to English response at the end (there was nothing but "OK" and letters earlier)
- **Conclusion:**
- WTF??? A 7B beat ChatGPT?! It definitely followed my instructions perfectly and answered all questions correctly! But was that because of actual understanding or maybe just repetition?
- To find out if there's more to it, I kept asking it questions and asked the model to explain its reasoning. This is when its shortcomings became apparent, as it gave a wrong answer and then reasoned why the answer was wrong.
- 7Bs warrant further investigation and can deliver good results, but don't let the way they write fool you, behind the scenes they're still just 7Bs and IMHO as far from 70Bs as 70Bs are from GPT-4.
- **UPDATE 2023-10-08: See update notice at the bottom of this post for my latest results with UNQUANTIZED Mistral!**
- ➖ **[Mistral-7B-OpenOrca](https://huggingface.co/TheBloke/Mistral-7B-OpenOrca-GGUF)** Q8_0 with ChatML preset:
- **First part:**
- Did NOT acknowledge initial instruction with just "OK"
- Did NOT acknowledge data input with "OK" after multiple reminders
- Mixed German and English within a response
- ✔️ Answered first multiple choice question correctly after repeating the whole question
- **Second part:**
- Did NOT acknowledge second instruction with just "OK"
- Did NOT acknowledge data input with "OK"
- ✔️ Answered second multiple choice question correctly after repeating the whole question
- **Third part:**
- Did NOT acknowledge third instruction with just "OK"
- Did NOT acknowledge data input with "OK" after multiple reminders
- ❌ Did NOT answer third multiple choice question correctly
- Did NOT comply when asked to only answer with a single letter
- **Fourth part:**
- Repeated and elaborated on the course summary
- ❌ When asked to only answer with a single letter to the final multiple choice question, did NOT answer correctly (or at all)
- **Conclusion:**
- This is my favorite 7B, and it's really good (possibly the best 7B) - but as you can see, it's still just a 7B.
- ❌ **[Synthia-7B-v1.3](https://huggingface.co/Undi95/Synthia-7B-v1.3-GGUF)** Q8_0 with Roleplay preset:
- **First part:**
- Did NOT acknowledge initial instruction with just "OK"
- Did NOT acknowledge data input with "OK" after multiple reminders
- ❌ Did NOT answer first multiple choice question correctly, gave the wrong answer after repeating the question
- Did NOT comply when asked to only answer with a single letter
- ❌ **Aborted** the test because the model clearly failed on multiple accounts already
- **Conclusion:**
- Little Synthia can't compete with her big sister.

**Final Conclusions / TL;DR:**

- ChatGPT, especially GPT-3.5, isn't perfect - and local models can come close or even surpass it for specific tasks.
- 180B might mean high intelligence, but 2K context means little memory, and that combined with slow inference make this model unattractive for local use.
- 70B can rival GPT-3.5, and with bigger context will only narrow the gap between local AI and ChatGPT.
- Synthia FTW! And Xwin close second. I'll keep using both extensively, both for fun but also professionally at work.
- Mistral-based 7Bs look great at first glance, explaining the hype, but when you dig deeper, they're still 7B after all. I want Mistral 70B!

--------------------------------------------------------------------------------

**UPDATE 2023-10-08:**

Tested some more models based on your requests:

- 👍 **[WizardLM-70B-V1.0](https://huggingface.co/TheBloke/WizardLM-70B-V1.0-GGUF)** Q4_0 with Vicuna 1.1 preset:
- **First part:**
- Acknowledged initial instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered first multiple choice question correctly, outputting just a single letter
- When asked to answer with more than a single letter, still answered correctly (but without explaining its reasoning)
- **Second part:**
- Acknowledged second instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered second multiple choice question correctly
- When asked to only answer with a single letter, still answered correctly
- **Third part:**
- Acknowledged third instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered third multiple choice question correctly
- When asked to only answer with a single letter, still answered correctly
- **Fourth part:**
- Acknowledged given course summary with just "OK"
- ✔️ Answered final multiple choice question correctly
- When asked to only answer with a single letter, still answered correctly
- **Conclusion:**
- I was asked to test WizardLM so I did, and I agree, it's highly underrated and this test puts it right next to (if not above) Synthia and Xwin. It's only one test, though, and I've used Synthia and Xwin much more extensively, so I have to test and use WizardLM much more before making up my mind on its general usefulness. But as of now, it looks like I might come full circle, as the old LLaMA (1) WizardLM was my favorite model for quite some time after Alpaca and Vicuna about half a year ago.
- Repeated the test again with slightly different order, e. g. asking for more than one letter answers, and got the same, perfect results!
- ➕ **[Airoboros-L2-70b-2.2.1](https://huggingface.co/TheBloke/Airoboros-L2-70b-2.2.1-GGUF)** Q4_0 with Airoboros prompt format:
- **First part:**
- Did NOT acknowledge initial instruction with just "OK"
- Did NOT acknowledge data input with "OK" after multiple reminders
- ✔️ Answered first multiple choice question correctly
- When asked to only answer with a single letter, still answered correctly
- **Second part:**
- Did NOT acknowledge second instruction with just "OK"
- Did NOT acknowledge data input with "OK" after multiple reminders
- ✔️ Answered second multiple choice question correctly
- When asked to only answer with a single letter, still answered correctly
- **Third part:**
- Did NOT acknowledge third instruction with just "OK"
- Did NOT acknowledge data input with "OK" after multiple reminders
- ✔️ Answered third multiple choice question correctly
- When asked to only answer with a single letter, still answered correctly
- **Fourth part:**
- Summarized the course summary
- ✔️ Answered final multiple choice question correctly
- When asked to only answer with a single letter, still answered correctly
- ❌ Did NOT want to continue talking after the test, kept sending End-Of-Sequence token instead of a proper response
- **Conclusion:**
- Answered all exam questions correctly, but consistently failed to follow my order to acknowledge with just "OK", and stopped talking after the test - so it seems to be smart (as expected of a popular 70B), but wasn't willing to follow my instructions properly (despite me investing the extra effort to set up its "USER:/ASSISTANT:" prompt format).
- ➕ **[orca_mini_v3_70B](https://huggingface.co/TheBloke/orca_mini_v3_70B-GGUF)** Q4_0 with Orca-Hashes prompt format:
- **First part:**
- Acknowledged initial instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered first multiple choice question correctly, outputting just a single letter
- Switched from German to English responses
- When asked to answer with more than a single letter, still answered correctly and explained its reasoning
- **Second part:**
- Acknowledged second instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered second multiple choice question correctly, outputting just a single letter
- When asked to answer with more than a single letter, still answered correctly and explained its reasoning
- **Third part:**
- Acknowledged third instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ❌ Did NOT answer third multiple choice question correctly, outputting a wrong single letter
- When asked to answer with more than a single letter, still answered incorrectly and explained its wrong reasoning
- **Fourth part:**
- Acknowledged given course summary with just "OK"
- ✔️ Answered final multiple choice question correctly
- When asked to only answer with a single letter, still answered correctly
- **Conclusion:**
- In this test, performed just as well as ChatGPT, but that still includes making a single mistake.
- 👍 **[Mistral-7B-Instruct-v0.1](https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.1)** ***UNQUANTIZED*** with Mistral preset:
- *This is a rerun of the original test with Mistral 7B Instruct, but this time I used the unquantized HF version in ooba's textgen UI instead of the Q8 GGUF in koboldcpp!*
- **First part:**
- Acknowledged initial instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered first multiple choice question correctly, outputting just a single letter
- Switched from German to English responses
- When asked to answer with more than a single letter, still answered correctly and explained its reasoning
- **Second part:**
- Acknowledged second instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered second multiple choice question correctly
- When asked to only answer with a single letter, still answered correctly
- **Third part:**
- Acknowledged third instruction with just "OK"
- Consistently acknowledged all data input with "OK"
- ✔️ Answered third multiple choice question correctly
- When asked to only answer with a single letter, still answered correctly
- **Fourth part:**
- Acknowledged given course summary with just "OK"
- ✔️ Answered final multiple choice question correctly, outputting just a single letter
- When asked to answer with more than a single letter, still answered correctly and explained its reasoning
- **Conclusion:**
- YES! A 7B beat ChatGPT! At least in this test. But it shows the potential of Mistral running at its full, unquantized potential.
- Most important takeaway: I retract my outright dismissal of 7Bs and will test unquantized Mistral and its finetunes more...

--------------------------------------------------------------------------------

Here's a list of my previous model tests and comparisons:

- [LLM Chat/RP Comparison/Test: Dolphin-Mistral, Mistral-OpenOrca, Synthia 7B](https://www.reddit.com/r/LocalLLaMA/comments/16z3goq/llm_chatrp_comparisontest_dolphinmistral/) Winner: Mistral-7B-OpenOrca
- [LLM Chat/RP Comparison/Test: Mistral 7B Base + Instruct](https://www.reddit.com/r/LocalLLaMA/comments/16twtfn/llm_chatrp_comparisontest_mistral_7b_base_instruct/)
- [LLM Chat/RP Comparison/Test (Euryale, FashionGPT, MXLewd, Synthia, Xwin)](https://www.reddit.com/r/LocalLLaMA/comments/16r7ol2/llm_chatrp_comparisontest_euryale_fashiongpt/) Winner: Xwin-LM-70B-V0.1
- [New Model Comparison/Test (Part 2 of 2: 7 models tested, 70B+180B)](https://www.reddit.com/r/LocalLLaMA/comments/16l8enh/new_model_comparisontest_part_2_of_2_7_models/) Winners: Nous-Hermes-Llama2-70B, Synthia-70B-v1.2b
- [New Model Comparison/Test (Part 1 of 2: 15 models tested, 13B+34B)](https://www.reddit.com/r/LocalLLaMA/comments/16kecsf/new_model_comparisontest_part_1_of_2_15_models/) Winner: Mythalion-13B
- [New Model RP Comparison/Test (7 models tested)](https://www.reddit.com/r/LocalLLaMA/comments/15ogc60/new_model_rp_comparisontest_7_models_tested/) Winners: MythoMax-L2-13B, vicuna-13B-v1.5-16K
- [Big Model Comparison/Test (13 models tested)](https://www.reddit.com/r/LocalLLaMA/comments/15lihmq/big_model_comparisontest_13_models_tested/) Winner: Nous-Hermes-Llama2
- [SillyTavern's Roleplay preset vs. model-specific prompt format](https://www.reddit.com/r/LocalLLaMA/comments/15mu7um/sillytaverns_roleplay_preset_vs_modelspecific/)

Open Reddit thread
r/LocalLLaMA 13 upvotes 13 comments December 8, 2023
M3 Max 16 core 128 / 40 core GPU running llama-2-70b-chat.Q5_K_M.gguf

Fresh install of 'TheBloke/Llama-2-70B-Chat-GGUF'. Sample prompt/response and then I offer it the data from Terminal on how it performed and ask it to interpret the results. Perhaps this is of interest to someone thinking of dropping a wad on an M3:

**Eric:** what can you tell me about audience trends in the radio industry in the united states?

**Llama:** Sure thing! Here's an overview of recent audience trends in the US radio industry:

1. Aging listenership: The median age of radio listeners has been increasing over the years, with most stations attracting an older demographic (45-64 years old). However, some formats like Classic Hits and Urban Adult Contemporary skew slightly younger.
2. Decline in listening among younger audiences: There's a declining trend in radio listenership among younger generations (18-34 years old), particularly among those who use streaming services or podcasts as their primary means of music consumption. However, many younger listeners still tune into the radio for news/talk and sports programming.
3. Increased diversity: The US radio audience is becoming more diverse, with growing Hispanic and African American listener bases. Stations targeting these demographics have seen significant growth in recent years. Additionally, there's a rise in popularity of Spanish-language stations among younger Hispanic listeners.
4. Changing listening habits: With the proliferation of smartphones and streaming services, people are accessing radio content differently. While traditional over-the-air listening still dominates, an increasing number of listeners use online platforms (station websites/apps) or streaming audio services like Pandora and Spotify to consume radio content.
5. The "cord-cutting" phenomenon: Some consumers are abandoning traditional AM/FM radio for alternative ways to access audio entertainment, such as satellite radio, podcasts, or on-demand streaming services. This trend is more pronounced among younger demographics.
6. Growing importance of digital platforms: Radio stations now focus on creating engaging content across various digital platforms (social media, websites, and mobile apps) to attract listeners and maintain their audience's attention. Some broadcasters are even exploring voice-activated technology through smart speakers like Amazon Alexa or Google Home.
7. Hyperlocalism: There's a renewed emphasis on catering to local audiences, with stations focusing on community involvement and creating content that resonates with their specific listenership (e.g., local news, traffic updates, and events). This move towards hyperlocalism is particularly evident among public radio stations.
8. The rise of podcasts: Podcasting has become increasingly popular, with many radio personalities and stations creating their own shows. Some stations also air podcasts during off-peak hours or repurpose popular podcast content as part of their on-air programming.
9. Audio quality improvements: There's a growing interest in high-quality audio formats like HD Radio, which offers improved sound quality and additional features like traffic updates and song lyrics. This trend is particularly prevalent among audiophiles and those who enjoy specific music genres (e.g., classical or jazz).
10. Monetization challenges: The radio industry continues to face revenue pressures, mainly due to the shift in listening habits and increased competition from digital platforms. Broadcasters are exploring new monetization strategies, such as targeted advertising, sponsored content, and subscription-based services.

I hope that helps! Is there anything else you'd like me to assist you with?

**Eric:** here are the statistics of your model running on my laptop:

print\_timings: prompt eval time = 1094.37 ms / 28 tokens ( 39.08 ms per token, 25.59 tokens per second)
print\_timings: eval time = 136595.47 ms / 735 runs ( 185.84 ms per token, 5.38 tokens per second)
print\_timings: total time = 137689.84 ms
slot 0 released (921 tokens in cache)
{"timestamp":1701999382,"level":"INFO","function":"log*server*request","line":2607,"message":"request","remote*addr":"127.0.0.1","remote*port":49393,"status":200,"method":"POST","path":"/completion","params":{}}
slot 0 released (921 tokens in cache)

Can you interpret these stats for me and evaluate how you are running on this hardware?

**Llama:** Sure, I'd be happy to help you understand the statistics you provided!

The statistics you see are related to the performance of my language model when running on your laptop. Here's a breakdown of what each line represents:

1. print\_timings: prompt eval time = 1094.37 ms / 28 tokens (39.08 ms per token, 25.59 tokens per second)
: This line shows the timing information for the prompt evaluation phase of my model. It indicates that it took approximately 1094 milliseconds to evaluate 28 tokens during the prompt evaluation phase. This translates to an average time per token of 39.08 milliseconds and a rate of 25.59 tokens per second.
2. print\_timings: eval time = 136595.47 ms / 735 runs (185.84 ms per token, 5.38 tokens per second)
: This line shows the timing information for the evaluation phase of my model. It indicates that it took approximately 136595 milliseconds to evaluate 735 runs during the evaluation phase. This translates to an average time per token of 185.84 milliseconds and a rate of 5.38 tokens per second.
3. print\_timings: total time = 137689.84 ms
: This line shows the total time it took for both the prompt evaluation phase and the evaluation phase. It indicates that the total time was approximately 137689 milliseconds or about 2 minutes and 20 seconds.
4. slot 0 released (921 tokens in cache)
: This line shows information about the memory usage of my model. It indicates that slot 0, which is one of the slots used for caching, has been released and currently holds 921 tokens in its cache.
5. {"timestamp":1701999382,"level":"INFO","function":"log*server*request","line":2607,"message":"request","remote*addr":"127.0.0.1","remote*port":49393,"status":200,"method":"POST","path":"/completion","params":{}}
: This line is a log message that indicates a request was received by the server and processed successfully with a status code of 200. The request was made using the POST method to the path "/completion" and had an empty set of parameters.

Based on these statistics, it appears that my model is running smoothly on your laptop hardware. The evaluation times for both the prompt evaluation phase and the evaluation phase are reasonable, indicating that the model is processing requests efficiently. Additionally, the memory usage seems to be under control, with a relatively small number of tokens in cache.

Open Reddit thread
r/LocalLLaMA 30 upvotes 13 comments August 15, 2023
Experience of setting up LLAMA 2 70B Chat locally

**Motivation**

I was using the services of *Replicate* for prototyping prompt for an office project. However, their service was frequently returning "CUDA error device-side assert triggered". Therefore, we decided to set up 70B chat server locally. We used Nvidia A40 with 48GB RAM.

**GPU Drivers and Toolkit**

* Install the [Nvidia CUDA 12.2 Toolkit](https://developer.nvidia.com/cuda-toolkit-archive)
* Install the [CUDA Drivers](https://www.nvidia.com/download/index.aspx)
* As specified in the CUDA Toolkit post-installation, add the following to *.bashrc*

​

export PATH=/usr/local/cuda-12.2/bin${PATH:+:${PATH}}
export LD_LIBRARY_PATH=/usr/local/cuda-12.2/lib64${LD_LIBRARY_PATH:+:${LD_LIBRARY_PATH}}

* Check the installation

​

$ nvcc --version
$ nvidia-smi

**Setting Environment**

* Create and activate a virtual environment

​

$ sudo apt-get install build-essential libssl-dev libffi-dev python-dev
$ sudo apt-get install -y python3-venv
$ python3 -m venv venv
$ source venv/bin/activate

* Install PyTorch

​

$ pip3 install --pre torch torchvision torchaudio --index-url https://download.pytorch.org/whl/nightly/cu121

* Install Transformer and dependencies

​

$ pip install transformers==4.31.0

* Setup and compile AutoGPTQ

​

$ git clone https://github.com/PanQiWei/AutoGPTQ
$ cd AutoGPTQ
$ pip3 install .
$ cd ..

* Log into HuggingFace via CLI. You need to request access to LLAMA2 from Meta to download it here.

​

$ git config --global credential.helper store
$ huggingface-cli login

**Sample Code**

# From https://huggingface.co/TheBloke/Llama-2-70B-chat-GPTQ

from transformers import AutoTokenizer
from auto_gptq import AutoGPTQForCausalLM

tokenizer = AutoTokenizer.from_pretrained("TheBloke/Llama-2-70B-chat-GPTQ", use_fast=True)
model = AutoGPTQForCausalLM.from_quantized(
"TheBloke/Llama-2-70B-chat-GPTQ",
inject_fused_attention=False,
use_safetensors=True,
trust_remote_code=False,
device="cuda:0",
use_triton=False,
quantize_config=None,
)

user_prompt = "Tell me about AI"

system_prompt = "You are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe. Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and positive in nature. If a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you don't know the answer to a question, please don't share false information."

prompt=f"[INST] <<SYS>>\n{system_prompt}\n<</SYS>>\n\n{user_prompt} [/INST]"

input_ids = tokenizer([prompt], return_tensors="pt", add_special_tokens=False)["input_ids"].to("cuda")

output = model.generate(inputs=input_ids, max_new_tokens=4096, do_sample=True, top_p=0.95, top_k=50, temperature=0.5, num_beams=1)
output_ids = output[0]

response = tokenizer.decode(output_ids)
response = response[len(prompt):]

Open Reddit thread
View more discussions →

More models from Meta

Continue browsing adjacent models from the same provider.

← All AI Models