NovatasticRoScript commited on
Commit
e2ac467
·
verified ·
1 Parent(s): 93b2fa6

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +83 -47
README.md CHANGED
@@ -1,59 +1,95 @@
1
  ---
 
2
  base_model: NovatasticRoScript/Atomight-2-1.5B-Thinking
3
  tags:
4
  - text-generation-inference
5
  - transformers
6
  - unsloth
7
- - qwen2
8
- - llama-cpp
9
- - gguf-my-repo
10
- license: apache-2.0
 
 
11
  language:
12
  - en
 
13
  datasets:
14
  - open-thoughts/OpenThoughts-114k
15
  ---
16
 
17
- # NovatasticRoScript/Atomight-2-1.5B-Thinking-Q4_K_M-GGUF
18
- This model was converted to GGUF format from [`NovatasticRoScript/Atomight-2-1.5B-Thinking`](https://huggingface.co/NovatasticRoScript/Atomight-2-1.5B-Thinking) using llama.cpp via the ggml.ai's [GGUF-my-repo](https://huggingface.co/spaces/ggml-org/gguf-my-repo) space.
19
- Refer to the [original model card](https://huggingface.co/NovatasticRoScript/Atomight-2-1.5B-Thinking) for more details on the model.
20
-
21
- ## Use with llama.cpp
22
- Install llama.cpp through brew (works on Mac and Linux)
23
-
24
- ```bash
25
- brew install llama.cpp
26
-
27
- ```
28
- Invoke the llama.cpp server or the CLI.
29
-
30
- ### CLI:
31
- ```bash
32
- llama-cli --hf-repo NovatasticRoScript/Atomight-2-1.5B-Thinking-Q4_K_M-GGUF --hf-file atomight-2-1.5b-thinking-q4_k_m.gguf -p "The meaning to life and the universe is"
33
- ```
34
-
35
- ### Server:
36
- ```bash
37
- llama-server --hf-repo NovatasticRoScript/Atomight-2-1.5B-Thinking-Q4_K_M-GGUF --hf-file atomight-2-1.5b-thinking-q4_k_m.gguf -c 2048
38
- ```
39
-
40
- Note: You can also use this checkpoint directly through the [usage steps](https://github.com/ggerganov/llama.cpp?tab=readme-ov-file#usage) listed in the Llama.cpp repo as well.
41
-
42
- Step 1: Clone llama.cpp from GitHub.
43
- ```
44
- git clone https://github.com/ggerganov/llama.cpp
45
- ```
46
-
47
- Step 2: Move into the llama.cpp folder and build it with `LLAMA_CURL=1` flag along with other hardware-specific flags (for ex: LLAMA_CUDA=1 for Nvidia GPUs on Linux).
48
- ```
49
- cd llama.cpp && LLAMA_CURL=1 make
50
- ```
51
-
52
- Step 3: Run inference through the main binary.
53
- ```
54
- ./llama-cli --hf-repo NovatasticRoScript/Atomight-2-1.5B-Thinking-Q4_K_M-GGUF --hf-file atomight-2-1.5b-thinking-q4_k_m.gguf -p "The meaning to life and the universe is"
55
- ```
56
- or
57
- ```
58
- ./llama-server --hf-repo NovatasticRoScript/Atomight-2-1.5B-Thinking-Q4_K_M-GGUF --hf-file atomight-2-1.5b-thinking-q4_k_m.gguf -c 2048
59
- ```
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ license: mit
3
  base_model: NovatasticRoScript/Atomight-2-1.5B-Thinking
4
  tags:
5
  - text-generation-inference
6
  - transformers
7
  - unsloth
8
+ - reasoning
9
+ - thought
10
+ - core-math
11
+ - instruction-tuning
12
+ model_creator: NovatasticRoScript
13
+ model_type: causal-lm
14
  language:
15
  - en
16
+ pipeline_tag: text-generation
17
  datasets:
18
  - open-thoughts/OpenThoughts-114k
19
  ---
20
 
21
+ <div align="center">
22
+
23
+ # ⚛️ Atomight-2-1.5B-Thinking
24
+
25
+ **A Deep-Reasoning Small Language Model Optimized for Sequential Logic Chains**
26
+
27
+ </div>
28
+
29
+ ## 📌 Model Overview
30
+ **Atomight-2-1.5B-Thinking** is a specialized, compact reasoning model built on top of a 1.5B parameter core architecture. Engineered explicitly for users operating on constrained hardware environments (such as a free Google Colab T4 instance), Atomight-2 utilizes an explicit internal `<think>...</think>` scratchpad layout. It dynamically breaks down complex mathematical, logical, and structural prompts before committing to a final conclusion.
31
+
32
+ ### 🚀 Key Highlights
33
+ * **Hardware Democratic:** High-tier deep reasoning accessible on consumer-grade hardware and free cloud compute tiers.
34
+ * **Structured Scratchpad:** Generates native, visible reasoning pathways natively formatted for transparent auditing.
35
+ * **Chat-Template Native:** Tailored directly for ChatML system configurations.
36
+
37
+ ---
38
+
39
+ ## 📊 Evaluation & Benchmark Results
40
+
41
+ Atomight-2 was subjected to a high-volume statistical evaluation matrix across core logic paradigms, matching up against premier industry baselines in the 1B–4B small language model class.
42
+
43
+ ### Official Performance Breakdown
44
+ The model displays exceptional specialization spikes in structured mathematical deduction, rivaling or outperforming significantly larger parameters classes on core numerical strings.
45
+
46
+ <div align="center">
47
+ <img src="https://huggingface.co/NovatasticRoScript/Atomight-2-1.5B-Thinking/resolve/main/Note%20Original%20benchmarking%20of%20Atomight-2-1.5B-Thinking%20consists%20of.png" alt="Atomight-2 Official Benchmark Result" width="85%">
48
+ </div>
49
+
50
+ | Benchmark | Paradigm | Atomight-2-1.5B-Thinking | Qwen-2-1.5B-Instruct | Phi-3-mini (3.8B) | Llama-3.2-3B-Instruct |
51
+ | :--- | :--- | :---: | :---: | :---: | :---: |
52
+ | **GSM8k** | Math Logical Chains | **80.1%** | 71.0% | 82.5% | 73.1% |
53
+ | **ARC-C** | Core Reasoning | **88.5%** | 82.3% | 84.9% | 83.3% |
54
+ | **MMLU** | General Knowledge | **63.2%** | 56.7% | 68.8% | 61.1% |
55
+
56
+ > ⚠️ **Evaluation Insight:** While Atomight-2 exhibits class-leading spikes on core textual logic and mathematical proofs, it experiences a classic reasoning tradeoff. On abstract matrix-grid visual transformation evaluations (like ARC-AGI 2), it drops to a baseline floor of **0.00%**. This cognitive bottleneck highlights an instruction deficit in translating spatial imagery into basic structural text tokens—a major priority slated for the next architecture generation.
57
+
58
+ ---
59
+
60
+ ## 💻 Quickstart & Inference Code
61
+
62
+ To deploy Atomight-2 cleanly without encountering text-truncation errors inside the internal reasoning blocks, execute the generation using the official structured chat template format.
63
+
64
+ ```python
65
+ import torch
66
+ from transformers import AutoTokenizer, AutoModelForCausalLM
67
+
68
+ MODEL_ID = "NovatasticRoScript/Atomight-2-1.5B-Thinking"
69
+
70
+ tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
71
+ model = AutoModelForCausalLM.from_pretrained(
72
+ MODEL_ID,
73
+ torch_dtype=torch.float16,
74
+ device_map="auto",
75
+ trust_remote_code=True
76
+ )
77
+
78
+ # Structure conversational dialog into ChatML framework
79
+ messages = [
80
+ {"role": "user", "content": "A retailer buys shirts for $15 and sells them for $25. What is the total profit on 12 shirts?"}
81
+ ]
82
+
83
+ templated_input = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
84
+ inputs = tokenizer(templated_input, return_tensors="pt").to("cuda")
85
+
86
+ print("🧠 Generating Reasoning Sequence:")
87
+ outputs = model.generate(
88
+ **inputs,
89
+ max_new_tokens=768, # Plentiful headroom required for deep-thinking scratchpads
90
+ temperature=0.1,
91
+ do_sample=False,
92
+ pad_token_id=tokenizer.eos_token_id
93
+ )
94
+
95
+ print(tokenizer.decode(outputs[0], skip_special_tokens=False))