97 5 197

Csaba Kecskemeti PRO

csabakecskemeti

https://devquasar.com/

csabakecskemeti

AI & ML interests

None yet

Recent Activity

updated a model about 18 hours ago

DevQuasar/mistralai.Mistral-Large-3-675B-Instruct-2512-GGUF

updated a model about 21 hours ago

DevQuasar/nvidia.Nemotron-Content-Safety-Reasoning-4B-GGUF

updated a model about 21 hours ago

DevQuasar/nvidia.Nemotron-Content-Safety-Reasoning-4B-GGUF

View all activity

Organizations

posted an update 3 days ago

Post

1092

FYI: Mistral.Ministral-3 dequantizer FP8->BF16

https://github.com/csabakecskemeti/ministral-3_dequantizer_fp8-bf16

(The instruct model weights are in FP8)

replied to their post 4 days ago

I've used this:
https://huggingface.co/meituan/DeepSeek-R1-Channel-INT8/tree/main/inference

Hoped I can make it work on my CPU... :P

replied to their post 4 days ago

@ubergarm you might have the resources!? 😀

posted an update 4 days ago

Post

2000

Looking for some help to test an INT8 Deepseek 3.2:
SGLang supports Channel wise INT8 quants on CPUs with AMX instructions (Xeon 5 and above AFAIK)
https://lmsys.org/blog/2025-07-14-intel-xeon-optimization/

Currently uploading an INT8 version of Deepseek 3.2 Speciale:
DevQuasar/deepseek-ai.DeepSeek-V3.2-Speciale-Channel-INT8

I cannot test this I'm on AMD
"AssertionError: W8A8Int8LinearMethod on CPU requires that CPU has AMX support"
(I assumed it can fall back to some non optimized kernel but seems not)

If anyone with the required resources (Intel Xeon 5/6 + ~768-1TB ram) can help to test this that would be awesome.

If you have hints how to make this work on AMD Threadripper 7000 Pro series please guide me.

Thanks all!

8 replies

posted an update 27 days ago

Post

301

Recently there are so much activity on token efficient formats, I've also build a package (inspired by toon).

Deep-TOON

My goal was to token efficiently handle json structures with complex embeddings.

So this is what I've built on the weekend. Feel free try:

https://pypi.org/project/deep-toon/0.1.0/

posted an update about 2 months ago

Post

2602

Christmas came early this year

3 replies

posted an update 6 months ago

Post

3067

Has anyone ever backed up a model to a sequential tape drive, or I'm the world first? :D
Just played around with my retro PC that has got a tape drive—did it just because I can.

5 replies

posted an update 6 months ago

Post

475

Deepseek R1 0528 Q2 locally.
(I believe it has overthinking it a bit :) )
https://youtu.be/Iqu5s9aFaXA?si=QWZe293iTKf_3ELU

DevQuasar/deepseek-ai.DeepSeek-R1-0528-GGUF

posted an update 8 months ago

Post

2119

Local Llama4 Maverick Q2
https://youtu.be/4F8g_LThli0?si=MGba2SUTHt6xYw3T
Quants uploading now

Big thanks to @ngxson !

posted an update 8 months ago

Post

1750

Why the 'how many r's in strawberry' prompt "breaks" llama4? :D

Quants DevQuasar/meta-llama.Llama-4-Scout-17B-16E-Instruct-GGUF

3 replies

posted an update 9 months ago

Post

3419

I'm collecting llama-bench results for inference with a llama 3.1 8B q4 and q8 reference models on varoius GPUs. The results are average of 5 executions.
The system varies (different motherboard and CPU ... but that probably that has little effect on the inference performance).

https://devquasar.com/gpu-gguf-inference-comparison/
the exact models user are in the page

I'd welcome results from other GPUs is you have access do anything else you've need in the post. Hopefully this is useful information everyone.

posted an update 9 months ago

Post

2412

Managed to get my hands on a 5090FE, it's beefy

| llama 8B Q8_0 | 7.95 GiB | 8.03 B | CUDA | 99 | pp512 | 12207.44 ± 481.67 |
| llama 8B Q8_0 | 7.95 GiB | 8.03 B | CUDA | 99 | tg128 | 143.18 ± 0.18 |

Comparison with others GPUs
http://devquasar.com/gpu-gguf-inference-comparison/

replied to their post 9 months ago

Follow-up

With the smaller context length dataset the training has succeeded.

posted an update 9 months ago

Post

1851

GTC new model announcement now from Nvidia
nvidia/Llama-3_3-Nemotron-Super-49B-v1

GGUFs:
DevQuasar/nvidia.Llama-3_3-Nemotron-Super-49B-v1-GGUF

Enjoy!

reacted to clem's post with 🚀 9 months ago

Post

4785

We just crossed 1,500,000 public models on Hugging Face (and 500k spaces, 330k datasets, 50k papers). One new repository is created every 15 seconds. Congratulations all!

3 replies

posted an update 9 months ago

Post

597

Cohere Command-a Q2 quant
DevQuasar/CohereForAI.c4ai-command-a-03-2025-GGUF

6.7t/s on a 3gpu setup (4080 + 2x3090)

(q3, q4 currently uploading)

replied to their post 9 months ago

No success so far, the training data contains some larger contexts and it fails just before complete the first epoch.
(dataset: DevQuasar/brainstorm-v3.1_vicnua_1k)

If anyone has further suggestion to the bnb config (with ROCm on MI100)?
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.bfloat16
)

Now testing with my other dataset that is smaller seems I have a lower memory need
DevQuasar/brainstorm_vicuna_1k

replied to their post 9 months ago

It's failed by the morning, need to find more room to decrease the memory

replied to their post 9 months ago

The machine itself is also funny. This my my GPU test bench.
Now also testing the PWM fan control and jetkvm

posted an update 9 months ago

Post

847

Fine tuning on the edge. Pushing the MI100 to it's limits.
QWQ-32B 4bit QLORA fine tuning
VRAM usage 31.498G/31.984G :D

4 replies

Csaba Kecskemeti PRO

AI & ML interests

Recent Activity

Organizations

csabakecskemeti's activity