Repository navigation
Expand file tree
/
Copy pathindex.html
More file actions
194 lines (179 loc) · 9.24 KB
/
Copy pathindex.html
File metadata and controls
194 lines (179 loc) · 9.24 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<meta
name="description"
content="Shuming Ma — research on language models, efficient architectures, reasoning, and multimodal systems."
/>
<meta name="author" content="Shuming Ma" />
<title>Shuming Ma</title>
<link rel="preconnect" href="https://fonts.googleapis.com" />
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin />
<link
href="https://fonts.googleapis.com/css2?family=IBM+Plex+Sans:wght@400;500;600;700&family=Source+Serif+4:opsz,wght@8..60,600;8..60,700&display=swap"
rel="stylesheet"
/>
<link rel="stylesheet" href="styles.css" />
</head>
<body>
<div class="page-shell">
<main class="page-main">
<section class="intro">
<div class="intro-copy">
<h1>Shuming Ma</h1>
<p class="name-alt">马树铭</p>
<p class="position">
Research on <strong>LLM pretraining</strong>, <strong>model architecture</strong>, and <strong>reasoning</strong>.
</p>
<p>
I work on large language models with an emphasis on scalable pretraining, efficient architectures, and
reasoning. Recent projects include <a href="https://arxiv.org/abs/2402.17764" target="_blank" rel="noreferrer">BitNet</a>,
<a href="https://aclanthology.org/2025.acl-long.457/" target="_blank" rel="noreferrer">bitnet.cpp</a>,
<a href="https://arxiv.org/abs/2407.10969" target="_blank" rel="noreferrer">Q-Sparse</a>,
<a href="https://arxiv.org/abs/2211.13184" target="_blank" rel="noreferrer">TorchScale</a>,
<a href="https://arxiv.org/abs/2307.02486" target="_blank" rel="noreferrer">LongNet</a>, and
<a href="https://arxiv.org/abs/2203.00555" target="_blank" rel="noreferrer">DeepNet</a>.
</p>
<p class="links-inline">
<a href="https://scholar.google.com/citations?user=J44tjDMAAAAJ&hl=en" target="_blank" rel="noreferrer">Google Scholar</a>
<span>/</span>
<a href="https://github.com/shumingma" target="_blank" rel="noreferrer">GitHub</a>
<span>/</span>
<a href="https://openreview.net/profile?id=~Shuming_Ma1" target="_blank" rel="noreferrer">OpenReview</a>
<span>/</span>
<a href="https://x.com/ma_shuming" target="_blank" rel="noreferrer">X / Twitter</a>
</p>
</div>
<div class="intro-photo">
<img src="assets/photo.jpg" alt="Portrait of Shuming Ma" />
</div>
</section>
<section class="section" id="news">
<h2>News</h2>
<div class="news-list">
<div class="news-item">
<div class="news-date">2025</div>
<div class="news-copy">
Introduced <strong>LongReasonArena</strong>, a benchmark for long reasoning that scales tasks to as much
as 1 million reasoning tokens.
<span class="item-links">
<a href="https://arxiv.org/abs/2508.19363" target="_blank" rel="noreferrer">[paper]</a>
</span>
</div>
</div>
<div class="news-item">
<div class="news-date">2025</div>
<div class="news-copy">
Released the <strong>BitNet b1.58 2B4T</strong> technical report for an open-source native 1-bit LLM at
the 2B scale trained on 4 trillion tokens.
<span class="item-links">
<a href="https://arxiv.org/abs/2504.12285" target="_blank" rel="noreferrer">[tech report]</a>
<a href="https://huggingface.co/microsoft/bitnet-b1.58-2B-4T" target="_blank" rel="noreferrer">[huggingface]</a>
</span>
</div>
</div>
<div class="news-item">
<div class="news-date">2025</div>
<div class="news-copy">
Published <strong>Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning</strong> on
more efficient test-time scaling.
<span class="item-links">
<a href="https://openreview.net/forum?id=6ICFqmixlS" target="_blank" rel="noreferrer">[paper]</a>
</span>
</div>
</div>
<div class="news-item">
<div class="news-date">2025</div>
<div class="news-copy">
Published <strong>bitnet.cpp</strong>, an inference stack for ternary and 1-bit LLMs aimed at efficient
edge inference.
<span class="item-links">
<a href="https://aclanthology.org/2025.acl-long.457/" target="_blank" rel="noreferrer">[paper]</a>
<a href="https://github.com/microsoft/BitNet" target="_blank" rel="noreferrer">[github]</a>
</span>
</div>
</div>
</div>
</section>
<section class="section" id="selected">
<h2>Selected Publications</h2>
<p class="section-note">
<em>For a full publication list, see Google Scholar.</em>
</p>
<div class="publication">
<div class="pub-title">The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits</div>
<div class="pub-venue">2024</div>
<div class="pub-summary">Introduces BitNet b1.58, showing that ternary 1-bit transformers can match full-precision baselines with substantially lower cost.</div>
<div class="pub-links">
<a href="https://arxiv.org/abs/2402.17764" target="_blank" rel="noreferrer">Paper</a>
</div>
</div>
<div class="publication">
<div class="pub-title">BitNet: Scaling 1-bit Transformers for Large Language Models</div>
<div class="pub-venue">2023</div>
<div class="pub-summary">Introduces BitNet, a scalable and stable 1-bit transformer architecture for large language model pretraining.</div>
<div class="pub-links">
<a href="https://arxiv.org/abs/2310.11453" target="_blank" rel="noreferrer">Paper</a>
</div>
</div>
<div class="publication">
<div class="pub-title">BitNet b1.58 2B4T Technical Report</div>
<div class="pub-venue">2025</div>
<div class="pub-summary">Presents the open-source 2B native 1-bit LLM trained on 4 trillion tokens and released with model weights.</div>
<div class="pub-links">
<a href="https://arxiv.org/abs/2504.12285" target="_blank" rel="noreferrer">Tech report</a>
<a href="https://huggingface.co/microsoft/bitnet-b1.58-2B-4T" target="_blank" rel="noreferrer">Hugging Face</a>
</div>
</div>
<div class="publication">
<div class="pub-title">bitnet.cpp: Efficient Edge Inference for Ternary LLMs</div>
<div class="pub-venue">2025</div>
<div class="pub-summary">An inference system for ternary and 1-bit LLMs with optimized kernels for efficient, lossless edge deployment.</div>
<div class="pub-links">
<a href="https://aclanthology.org/2025.acl-long.457/" target="_blank" rel="noreferrer">Paper</a>
<a href="https://github.com/microsoft/BitNet" target="_blank" rel="noreferrer">GitHub</a>
</div>
</div>
<div class="publication">
<div class="pub-title">Q-Sparse</div>
<div class="pub-venue">2024</div>
<div class="pub-summary">Training LLMs with fully sparsely-activated linear transformations for more efficient inference.</div>
<div class="pub-links">
<a href="https://arxiv.org/abs/2407.10969" target="_blank" rel="noreferrer">Paper</a>
</div>
</div>
<div class="publication">
<div class="pub-title">TorchScale: Transformers at Scale</div>
<div class="pub-venue">2022</div>
<div class="pub-summary">An open-source toolkit for scaling transformers, including architectures such as DeepNet and LongNet.</div>
<div class="pub-links">
<a href="https://arxiv.org/abs/2211.13184" target="_blank" rel="noreferrer">Paper</a>
<a href="https://github.com/microsoft/torchscale" target="_blank" rel="noreferrer">GitHub</a>
</div>
</div>
<div class="publication">
<div class="pub-title">LongNet: Scaling Transformers to 1,000,000,000 Tokens</div>
<div class="pub-venue">2023</div>
<div class="pub-summary">Introduces dilated attention to scale transformer context length to more than 1 billion tokens.</div>
<div class="pub-links">
<a href="https://arxiv.org/abs/2307.02486" target="_blank" rel="noreferrer">Paper</a>
</div>
</div>
<div class="publication">
<div class="pub-title">DeepNet: Scaling Transformers to 1,000 Layers</div>
<div class="pub-venue">2022</div>
<div class="pub-summary">Introduces DeepNorm and initialization strategies that stabilize extremely deep transformer training.</div>
<div class="pub-links">
<a href="https://arxiv.org/abs/2203.00555" target="_blank" rel="noreferrer">Paper</a>
</div>
</div>
</section>
</main>
<footer class="site-footer">
<p>Updated April 2026.</p>
</footer>
</div>
</body>
</html>