Math journal suggester

Methodology

An experiment in recommending math journals from a paper’s title and abstract using Qwen3-Embedding-8B and Kev-4B

How recommendations work

Qwen3-Embedding-8B embeds the title and abstract, then finds similar published papers by cosine similarity. Each journal is scored by its closest paper. The website currently shows the five highest-scoring journals, with up to two real supporting papers each. The 20 closest journals and their supporting excerpts are then passed to a fine-tuned Kev-4B for reranking.

Data collection

The audited corpus covers 95 journals and series from 2016 through 2025. It contains 10,061 reference papers, plus separate sets of 1,000 validation and 1,000 test papers. Training papers come from the reference pool.

Publication metadata comes from Crossref and publisher records; available abstracts come from Crossref, OpenAlex, publishers, and arXiv, with identity checks before matching. Years refer to journal publication, not arXiv posting.

Sampling approximately follows indexed publication counts; validation and test each include at least one paper per journal.

Results so far

Top 3 accuracy is the fraction of papers whose actual publication journal appears among the first three suggestions. All methods use the same papers and natural candidate lists; missing the actual journal counts as a failure.

MethodTop 1Top 3
Qwen retrieval21.8%41.7%
Qwen + released Kev21.6%39.2%
Qwen + tuned Kev28.1%50.8%

The actual journal reaches the top 20 candidate list for 84.5% of validation papers. Kev cannot recover the remaining 15.5%. The tuned result was selected using this validation set, so it is not an independent test result.

Final benchmark: after selecting the checkpoint, evaluate these three methods once on the untouched 1,000-paper test set. Report top 1, top 3 and a paired 95% confidence interval for the gain over retrieval.

Learning-rate comparison

Each trial started from released Kev, using the same 1,000 training papers, seed and two-epoch budget. The backbone stayed frozen; the existing adapter and decision head were updated. Every epoch was evaluated on all 1,000 validation papers.

Top 3 validation accuracy by peak learning rate
Peak rateEpoch 1Epoch 2
1e-547.4%49.5%
2e-548.2%50.3%
4e-549.9%50.8%

We selected 4e-5 by validation top 3 accuracy. Its 0.5 percentage-point lead over 2e-5 is uncertain (descriptive paired 95% interval: −1.1 to +2.2 points). This small search does not establish an optimal rate.

Validation cross-entropy fell from 2.264 to 2.198, 2.210 to 2.182, and 2.167 to 2.162, respectively. Loss uses the fixed 845 papers whose target journal was retrieved; accuracy includes all 1,000. Completed trials took 16–20 minutes each, including validation.

8,000-paper training protocol · results pending

The expanded run starts again from released Kev with 8,000 training papers, proportional journal quotas and shortages redistributed. There is no minimum-ten floor; journal counts range from 6 to 399.

Peak learning rate 4e-5; AdamW with a five-epoch OneCycle schedule; microbatch one, accumulation eight; at most five epochs. Validate and save every half epoch. After at least two epochs, stop after three checks without a meaningful loss improvement (0.01) or a new best top 3 score. Small loss improvements accumulate. Select the highest validation top 3 checkpoint, preferring the earlier checkpoint on ties.

All 8,000 requests passed independent checks, with no rejected or truncated examples. For 1,191 training papers, the target journal was inserted because retrieval missed it; validation and test never receive this insertion. A three-update pilot verified gradients, saved state and identical predictions after reloading.

The original call was cancelled before its first validation checkpoint; the cause is unconfirmed. The identical run was restarted from released Kev with an independent cloud job and heartbeat monitoring. Final results will be added after export and audit.

What this measures

Matching the actual journal measures agreement with historical publication choices. Other journals may also be suitable; the scores are not acceptance probabilities. Available abstracts bias the sample, larger journals can benefit from having more reference papers, and these splits do not rule out exposure during model pretraining.