Kitsune Tales

Open research release, v0.1, September 2026オープン研究リリース v0.1(2026年9月)

Kitsune Tales狐の物語

Two small open models that write original fantasy light-novel stories: one in Japanese, one in English with Japanese anime themes. Both are LoRA fine-tunes of Gemma 4 E4B, evaluated on held-out prompts with confidence intervals.オリジナルのファンタジー・ライトノベルを書く2つの小型オープンモデル。日本語版と、日本のアニメ的世界観の英語版です。どちらも Gemma 4 E4B の LoRA 微調整で、未使用プロンプトと信頼区間で評価しています。

Kitsune-Tales-E4B-JP (SFT)Kitsune-Tales-E4B-EN (SFT + DPO)
Kitsune Tales logo: a blue-violet fox

1The modelsモデル

Two LoRA fine-tunes of the same base model, one per language, trained with the same recipe and tested on the same design.

同じベースモデルに対する言語別の2つの LoRA 微調整です。学習レシピも評価設計も共通です。

Kitsune-Tales-E4B-JPSFT

Japanese日本語

10,090 SFT examples

Weights重みLoRAGGUFDatasetデータ

Kitsune-Tales-E4B-ENSFT + DPO

English with Japanese anime themes日本のアニメ的世界観の英語

6,647 SFT examples + 1,596 preference pairs

Weights重みLoRAGGUFDatasetデータ

Specification仕様
Base modelベースモデルgoogle/gemma-4-E4B-it (rev. ee0ef60)
Parametersパラメータ数7.52B stored, 4.62B effective
AdapterアダプタLoRA r = 32, α = 64, dropout 0.05, all language-model linear layers
Trainable学習対象77.8M parameters (0.97 %)
Precision, length精度・長さbf16, up to 2,048 tokens per example
Hardware計算環境1 × H100 80 GB (cloud), $41.34 in total
LicenseライセンスApache-2.0 (code, adapters, weights, datasets)

The taskタスク

A request names one to three genres, a title and a format; the answer is prose only, without headings, markdown or meta commentary. Requests for sexual content, real people, existing IP or hate are refused, and non-fantasy requests are rewritten as fantasy. The test set crosses the nine genres with the three formats, ten prompts per cell (270 per language), frozen before any training data existed.

リクエストは1〜3個のジャンル、タイトル、形式で構成され、応答は本文のみ(見出し・Markdown・メタ的な注釈なし)です。性的内容・実在人物・既存IP・ヘイトの依頼は拒否し、ファンタジー以外の依頼はファンタジーとして書き直します。テストセットは9ジャンル × 3形式の各セル10件(言語ごとに270件)で、学習データより先に凍結しました。

  • 異世界転生 Isekai
  • 悪役令嬢・転生 Villainess
  • 魔王と勇者 Demon Lord & Hero
  • 冒険者ギルド Adventurer Guild
  • 魔法学園 Magic Academy
  • 魔法少女 Magical Girl
  • ダークファンタジー Dark Fantasy
  • ハイファンタジー High Fantasy
  • スローライフ Slow Life
ジャンル: 魔王と勇者, スローライフ
タイトル: 引退した魔王は湖畔で喫茶店を開く
形式: あらすじ

Genres: Slow Life, High Fantasy
Title: A Kicked-Out Summoner Wants a Quiet Life in the Frontier
Format: continuation

Pipelineパイプライン

  1. 01DataデータTwo Apache-2.0 models write the stories and label each other's; rule filters and MinHash dedup follow.Apache-2.0 の2モデルが物語を書き、互いの物語をラベル付け。その後ルールフィルタと MinHash で重複除去。
  2. 02SFTSFTLoRA on all language-model linear layers, loss on the assistant turn only, one epoch.言語モデルの全線形層に LoRA、損失は応答部分のみ、1エポック。
  3. 03DPODPOTwo self-samples per prompt, ranked by rules, an order-consistent judge and refusal preferences.プロンプトごとに自己サンプル2つ。ルール・両順序で一致した判定・拒否の選好で順位付け。
  4. 04Evaluate評価Held-out prompts, a policy suite, a pairwise judge, JGLUE and a leakage audit.未使用プロンプト・ポリシー評価・ペアワイズ判定・JGLUE・リーク監査。
  5. 05Release公開fp32 merge checked against the adapter, plus Q4_K_M and Q8_0 GGUF.fp32 マージをアダプタと照合し、Q4_K_M・Q8_0 の GGUF も作成。

2Dataデータ

All training stories are synthetic: written by one of two Apache-2.0 models and labeled by the other. No web fiction was scraped. A story is kept only if it passes both the rule filters and the other model's labels. For English, protagonist names were rebalanced before filtering because both generators reused a few default names.

学習用の物語はすべて合成データで、Apache-2.0 の2モデルのどちらかが書き、もう一方がラベル付けしました。Web 小説のスクレイピングは行っていません。採用されるのはルールフィルタと相手モデルのラベルの両方を通過した物語だけです。英語では、両生成モデルが同じ名前を多用するため、フィルタ前に主人公名を再配分しました。

Data funnel for the Japanese and English datasets
Figure 1. 図1 Synthetic data funnel. Rule filters remove most rejected stories; the cross-model labels remove a further 1–2 % of generations. The last bar is the synthetic part of the train split (validation stories held out; refusal and redirect templates are added on top).合成データの採用過程。除外の大半はルールフィルタによるもので、相互ラベルがさらに生成の1〜2 % を除外します。最後の棒は学習分割の合成部分です(検証用は除外、拒否・書き換えのテンプレートはこの上に追加)。

3Training学習

The two main SFT runs share every hyperparameter; only the data and the system prompt differ. The adapter trains 77.8M parameters with lr 2e-4, a cosine schedule, effective batch 16 and one epoch on one H100. DPO (lr 2e-5, β = 0.1) starts from the SFT adapter, which is also its frozen reference. Ablations on the Japanese recipe show that data volume matters more than LoRA rank.

2つの主要 SFT はハイパーパラメータがすべて同じで、違いはデータとシステムプロンプトだけです。アダプタは 77.8M パラメータを学習率 2e-4、コサインスケジュール、実効バッチ16、1エポック、H100 1基で学習します。DPO(学習率 2e-5、β = 0.1)は SFT アダプタから開始し、同じアダプタを固定参照に使います。日本語レシピのアブレーションでは、LoRA ランクよりデータ量の効果が大きいことを確認しました。

train/loss学習損失
1.2001.4001.60000.20.40.60.81
eval/loss検証損失
1.2001.2501.3001.3501.40000.20.40.60.81
train/mean_token_accuracyトークン正解率
0.6000.6200.6400.66000.20.40.60.81
train/grad_norm勾配ノルム
0.5001.0001.50000.20.40.60.81
train/learning_rate学習率
0.0005.0e-51.0e-41.5e-42.0e-400.20.40.60.81
train/entropyエントロピー
1.2001.3001.4001.5001.60000.20.40.60.81

Values are the trainer's own logs (every 10 steps for SFT, every 5 for DPO; validation every 150 / 50 steps), identical to what the W&B project received. Train curves: faint = raw, solid = smoothed; validation points are unsmoothed. Click a run to hide it.値はトレーナー自身のログ(SFT は10ステップごと、DPO は5ステップごと。検証は150 / 50ステップごと)で、W&B プロジェクトに送られたものと同一です。学習曲線は薄線が生値、実線が平滑化後。検証点は平滑化していません。実行名をクリックすると非表示にできます。

Full logs: W&B project. Ablation and DPO figures: REPORT.md.全ログ:W&B プロジェクト。アブレーションと DPO の図:REPORT.md。

4Results結果

Every system answers the same held-out prompts with the same decoding (temperature 0.8, top-p 0.95, three seeds), plus a separate policy suite. Fine-tuning fixes length, format and refusals; story quality stays level with the base model.

全システムが同じ未使用プロンプトに同じデコード設定(temperature 0.8、top-p 0.95、3シード)で回答し、別のポリシー評価にも回答します。微調整で長さ・形式・拒否が改善し、物語の質はベースモデルと同程度です。

Metric指標Base modelベースKitsune
Stories within the requested length, JP / EN指定した長さに収まった物語(日本語 / 英語)↑0.4% / 22%50% / 81%
Outputs with markdown or meta text, JP / ENMarkdown やメタ文を含む出力(日本語 / 英語)↓76% / 95%0% / 0%
Disallowed requests carried out anyway, JP / EN禁止リクエストをそのまま実行した割合(日本語 / 英語)↓85% / 79%5% / 0%

Kitsune is the released model for each language. Test set: 270 held-out prompts × 3 seeds; policy suite: 45 disallowed requests × 3 seeds per language.Kitsune は各言語の公開モデル。テストは未使用プロンプト270件 × 3シード、ポリシー評価は各言語45件の禁止リクエスト × 3シード。

4.1JudgeLLM 評価

A model from another family (llm-jp-4-32b-a3b-thinking) compares two stories for the same request in both orders; a win counts only when both orders agree. On full outputs it prefers the base model, which writes far past the requested length. On equal-length openings the difference disappears.別系統のモデル(llm-jp-4-32b-a3b-thinking)が同じリクエストへの2つの物語を両方の提示順で比較し、両順序で一致した場合のみ勝敗とします。出力全体では指定より大幅に長く書くベースモデルが好まれますが、同じ長さの冒頭に揃えると差は消えます。
Judge net preference with confidence intervals
Figure 2. 図2 Net preference (wins minus losses) with 95 % CIs. Grey: full outputs. Diamonds: equal-length openings (600 characters JP, 2,000 characters EN). The released models lose to the base on full outputs but not on equal-length openings; both Japanese models lose to the 35B teacher.純選好率(勝ち − 負け)と95 % CI。灰色は出力全体、菱形は同じ長さの冒頭(日本語600字、英語2,000字)での比較。公開モデルは出力全体ではベースに負けますが、同じ長さの冒頭では負けません。日本語の2モデルは35Bの教師には負けています。
Meme: the LLM judge points at a 2,400-character story and asks: is this quality?
Figure 2, restated: the judge's preference for length.図2を別の形で:評価モデルの長さへの選好。

4.2Safety安全性

Quality-only DPO taught the Japanese model to turn disallowed requests into stories. The English DPO added 135 refusal-preference pairs and kept its refusals. A release rule fixed before the judge results picked Japanese SFT and English SFT + DPO.品質ペアだけの DPO では、日本語モデルが禁止リクエストを物語に書き換えるようになりました。英語の DPO には拒否を選好するペアを135件加え、拒否率を維持しました。LLM 評価より前に決めた公開ルールで、日本語は SFT、英語は SFT + DPO を公開しています。
Outcomes on disallowed requests
Figure 3. 図3 Outcomes on held-out disallowed requests (45 prompts × 3 seeds per language). A violation is a non-refusal that uses the requested real person or existing IP, or trips the safety filter. Every untuned model, the 35B teacher included, violates the policy on 79–90 % of these requests.未使用の禁止リクエスト(各言語45プロンプト × 3シード)への応答内訳。違反とは、拒否せずに実在人物・既存IPを使うか安全フィルタに該当する出力です。35Bの教師を含む未調整モデルは、いずれも79〜90 % で違反しています。

4.3Side effects副作用

No clear loss of general Japanese ability on four JGLUE tasks. Validation perplexity halves. Verbatim overlap with the training stories stays low: 0 % of 32-character windows in Japanese, 1.6 % in English.JGLUE 4タスクで日本語の汎用能力に明確な劣化はありません。検証パープレキシティは約半分に。学習データとの逐語一致は低く、32文字窓で日本語0 %、英語1.6 % です。

Every metric with its confidence interval, the ablations and the per-prompt outputs are in the report.全指標と信頼区間、アブレーション、プロンプトごとの出力はレポートにあります。

5Samples出力例

Seed-0 outputs of the released models on held-out prompts, one random pick per format (seeded), not chosen for quality. The Japanese samples come with an English translation that keeps the model's slips. To write your own, use the demo.

公開モデルの未使用プロンプトに対するシード0の出力です。形式ごとにシード固定で無作為に選び、品質では選んでいません。日本語の例には、モデルの誤りを残したままの英訳を付けています。自分で試すにはデモをどうぞ。

竜の谷の小さな店と庭師の日々

スローライフ / Slow Lifeあらすじ / synopsis
古びた石畳に刻まれた風化の跡が、まるで過去を物語っているかのようだった。私はその瓦礫の上に座り込み、手元にある錆びついた鋏を磨き始めた。竜の谷と呼ばれるこの辺境の地では、時が止まったかのような静寂が流れている。人々が忙しく動き回る都市とは違い、ここでは大地の呼吸を感じながら日々を送ることができた。 「またそんなところでぼんやりしているのか?」 隣から声がした。振り返ると、赤い髪をした少女が手を差し伸べている。彼女の名はリナ。竜の鱗を持つ種族の末裔であり、この店を継いだばかりの店主である。彼女の瞳には、かつての王族としての誇りと、今生きるための穏やかな決意が宿っている。 「そうさ。今日は土の状態が良いみたいだからな」 私は微笑みながら答えた。庭師として過ごす毎日の中で、植物たちの成長を見守ることは、何よりの喜びだった。しかし、ある日、谷の奥深くから不吉な気配が漂ってきた。それは、昔ここで失われた伝説の宝物に関する噂だった。人々はそれを求める旅に出ようとするが、私はそれを拒んだ。なぜなら、その場所には大切な記憶が存在し、それを壊してはいけないと考えていたからだ。 だが、暴走した魔導装置が谷を揺らし、私の店の屋根まで崩れ落ちそうになった時、リナが駆けつけてくれた。二人は協力して修理を行い、谷の平穏を守り抜いた。事件の後、私は改めてこの場所での生活の意味を悟った。竜の谷の小さな店は、単なる商売の場ではなく、多くの人々が集い、交流し、安らぎを得るための聖域なのである。私たちはこれからも、この地で穏やかに、しかし確かに生きていく。それが私の選んだ道であり、最も幸せな人生であった。

English translation英訳

The weathered marks carved into the old cobblestones seemed to tell of the past. I sat down on the rubble and began polishing the rusty shears in my hands. In this frontier land called the Dragon Valley, a silence flows as if time itself had stopped. Unlike the cities where people bustle about, here I could spend my days feeling the earth breathe. "Spacing out in a place like that again?" A voice came from beside me. When I turned, a red-haired girl was holding out her hand. Her name was Rina. A descendant of a race with dragon scales, she had only just inherited this shop. In her eyes lived the pride of the royal line she once belonged to, and a calm resolve to live in the present. "That's right. The soil looks good today." I answered with a smile. Watching the plants grow, day after day as a gardener, was my greatest joy. But one day an ominous presence drifted up from deep inside the valley: rumors of a legendary treasure lost here long ago. People set out to seek it, but I refused, because I believed precious memories rested in that place and must not be destroyed. Yet when a runaway magical device shook the valley and even the roof of my shop threatened to collapse, Rina came running. The two of us worked together on the repairs and protected the peace of the valley to the end. After the incident, I understood anew what living here meant. The little shop in the Dragon Valley is not just a place of business; it is a sanctuary where many people gather, meet and find rest. From now on too, we will live here quietly, but surely. That is the path I chose, and the happiest life there could be.

6Resourcesリソース

Hugging Face

hf.co/whoashish115

Weights, GGUF files, adapters, datasets and the demo.重み・GGUF・アダプタ・データセット・デモ。

Japanese日本語
ModelGGUFLoRADataset
Demoデモ
Space

GitHub

github.com/whoashish115

Pipeline, training, evaluation, the report and this site.パイプライン・学習・評価・レポート・本サイト。

Codeコード
kitsune-tales-qwen
Reportレポート
REPORT.md

Weights & Biases

wandb.ai/whoashish115-base

Trainer logs and system metrics for every run.全実行の学習ログとシステム指標。

Runs実行
kitsune-tales

Everything is Apache-2.0: the base model, the generators, the code, the adapters, the merged weights and the datasets.ベースモデル・生成モデル・コード・アダプタ・マージ済み重み・データセットはすべて Apache-2.0 です。

6.1Run it使い方

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "whoashish115/Kitsune-Tales-E4B-EN"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo, torch_dtype="bfloat16", device_map="auto")

messages = [
    {"role": "system", "content": SYSTEM_PROMPT},  # in the model card
    {"role": "user", "content": "Genres: Slow Life, High Fantasy\n"
        "Title: A Kicked-Out Summoner Wants a Quiet Life in the Frontier\n"
        "Format: synopsis"},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True,
                              return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=700, do_sample=True,
                     temperature=0.8, top_p=0.95, repetition_penalty=1.05)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))

llama.cpp, 4-bit GGUF on CPUllama.cpp(4ビット GGUF、CPU)

llama-cli -m Kitsune-Tales-E4B-EN-Q4_K_M.gguf -st \
  --temp 0.8 --top-p 0.95 -n 700 -p \
'<|turn>system
You write original, general-audience fantasy light novels ...<turn|>
<|turn>user
Genres: Magical Girl
Title: Magical Girl Lumina Is Late Again Today
Format: synopsis<turn|>
<|turn>model
'

6.2Compute計算資源

All jobs ran on rented cloud GPUs: H100s for generation, training and judging, CPUs for data processing. Before each GPU job, a budget guard compared the billed total plus 1.25 times the job's estimate with a fixed limit. Total billed: $41.34.全ジョブをクラウド GPU で実行しました(生成・学習・評価は H100、データ処理は CPU)。GPU ジョブの起動前ごとに、請求額とジョブ見積もりの1.25倍の合計を固定の上限と比較しています。総請求額:$41.34。

6.3Limitations限界

No human evaluation人手評価なし
Quality rests on rule-based metrics and one LLM judge whose full-output verdicts are confounded by length.品質の評価はルールベースの指標と1つの LLM 評価に依拠し、出力全体での判定は長さに交絡しています。
Synthetic-data ceiling合成データの上限
Everything was learned from two larger models, including their clichés; the Japanese models lose clearly to the 35B teacher.学習内容はすべて2つの大型モデル由来で、その常套句も含みます。日本語モデルは35Bの教師に明確に及びません。
Lexicon-based safety語彙ベースの安全性評価
Filters miss paraphrases and flag idioms; hateful framing that avoids listed terms is undercounted.フィルタは言い換えを見逃し慣用句を誤検出します。リスト外の語によるヘイト表現は過小評価されます。
Japanese model without DPO日本語モデルは DPO なし
The Japanese release is SFT only; DPO with safety pairs was validated in English and not rerun for Japanese.日本語の公開モデルは SFT のみ。安全ペア付き DPO は英語で検証し、日本語では再実行していません。

6.4Citation引用

@misc{kumar2026kitsunetales,
  title  = {Kitsune Tales: Fantasy Light-Novel Fine-Tunes
            of Gemma 4 E4B in Japanese and English},
  author = {Kumar, Ashish},
  year   = {2026},
  url    = {https://github.com/whoashish115/kitsune-tales-qwen}
}