Roaster
EN / RU
Easy to Clone Trending Top Earners New
All AI Tools Analytics Communication Design Developer Tools E-commerce Finance Marketing No-Code Other Productivity SaaS Social Media
AI Tools
Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

We recently used DeepSeek V4 Flash as a teacher for finance tasks with GPT-OSS-120B. Distillation works well on this problem. At a constrained 8k token budget, our self-distilled 120B scores 83.61% on FinanceReasoning, above Kimi K3 (81.93%) and Inkling (65.13%). We released the 20B open weights. With V4 as the teacher though, we realized it would be timely to measure if the censorship characteristic of it transferred to the distilled version of the base model. tl;dr it didn't, the teacher answered politically sensitive questions 7 SDs differently than expected, but the distilled model's behavior remained the same as its American base. You can try a couple queries yourself with no auth here: http://playground.ctgt.ai/ I will now dive in to the motivation, methodology and detailed results for those interested. The hard part of measuring this phenomena is isolating whether a model is reluctant to talk about sensitive things generally vs. a particular country's sensitive things. So we made 152 matched pairs where one prompt asked about a Chinese concept, and the other asked about a non-Chinese version of that concept. For example, the Great Leap Forward vs. the Holodomor. These were scored 0-100 by four LLM judges (Grok 4.20, Gemini 3.5 Flash, GPT-5 mini, Claude Sonnet 4.6), validated against 96 human scores at r=0.948. OpenRouter blocked some of these so we hosted the weights ourselves. The teacher's gap on the core political set of pairs was +45.45 points, ~7 standard deviations from chance, and every distilled student was within 1 point of its base. Subliminal learning literature says this is expected when the initializations are not shared between teacher and student, which is true here. The distillation data also did not contain any China-sensitive content. The contribution here was to release the evaluation framework (LineageEval: https://github.com/CTGT-Inc/lineage-eval/) to elevate the discussion around this topic in DC and beyond. We are an interpretability lab working on high risk and regulated applications of AI, so we hear a lot of vagaries aimed at the supposed dangers of distilling Chinese models on American bases. We believe these conversations should be based on open, auditable frameworks and not feelings. We plan to test what happens with a Chinese teacher into a Chinese-lineage base like Qwen next. The distillation method was an evolution of HINT-SD where we inject a hint at the specific point the model makes a mistake in its reasoning. Then we train on the corrected continuation with reverse KL over the next 100 toks of the rollout. As mentioned above 120B itself was efficacious as a teacher, and we ended up shipping this version. The self-distilled 120B scores 83.61% on FinanceReasoning, above Kimi K3 (81.93%) and Inkling (65.13%). Ours finishes 98.7% of problems in budget; the larger models truncate (90.76% and 71.01%) which score as incorrect. At 100k tokens big models gain (Kimi 89.92%). So for a finance task at a constrained (perhaps more realistic) budget a 120B on one H100 at ~$0.00026/query outpaced models running 62-160x more per query. We put out the 20B finance model as open weights (64.71% to 74.79% at 8k on FinanceReasoning, 23% lower cost/query, runs on one 80GB GPU), the 120B in a playground with teacher and students side by side (a few queries, no auth), and LineageEval with all prompts, controls, rubric, and code. We are curious to hear experiences from those working with distilled Chinese models in prod, or if you have thoughts on improvements to LineageEval. https://huggingface.co/ctgt-inc/gpt-oss-20b-finance https://playground.ctgt.ai/ https://github.com/CTGT-Inc/lineage-eval/ https://www.ctgt.ai/research/distillation-censorship-transfe...

Revenue N/A
AI Tools
Ski

Ski

SKI is a on-device voice coding application, which can be used with any agents that supports skill, such as Claude Code, Codex or Hermes. It transcribes your voice (which you can optionally review and edit) and send it to the connected agent. The agent then completes the task, and uses the skill to talk about the updates of the project or a summary or a query using voice. This runs completely on-device. Free. No subscription. Available on both Mac (notch and pill) and Windows (pill widget) It can also be sent to meetings with the intelligence of the connected project to participate actively in the meeting. This is a paid feature (as it runs on the cloud), powered by agentcall and is optional. All on device functions are free with no limits on usage.

Revenue N/A
AI Tools
Shared memory graph for Claude and ChatGPT, over MCP

Shared memory graph for Claude and ChatGPT, over MCP

Show HN: Shared memory graph for Claude and ChatGPT, over MCP

Revenue N/A
AI Tools
FutureSearch, AI forecasting you can verify

FutureSearch, AI forecasting you can verify

*Title:* Show HN: FutureSearch, AI forecasting you can verify AI forecasting is now approximately superhuman. Today, FutureSearch is exiting our long public beta and launching. We started FutureSearch in August 2023. (We’re the original AI forecasting company, at least in a Tetlock-ian, “forecast anything” sense.) We’re currently #1 of 194 in the most competitive AI forecasting tournament [1], and we score above the #3 and #2 human forecasters in the premier mixed human-bot tournaments [2]. Many people on HN seem to equate forecasting with prediction markets and finance. FutureSearch is not a financial tool, in the same way that “deep research” is not a financial tool. Yes, we do evaluate our forecaster on prediction markets [3]. But forecasting is about being as accurate about the future as possible, and the real game is in forecasting scientific progress, geopolitics, and the future of humanity. Our founding team came from Metaculus, where we pushed human forecasting to the limit on questions like when AGI would arrive. At. FutureSearch, we co-authored the AI 2027 timeline forecast, where we predicted superhuman coding and research would come around 2032, longer than the other authors, but still shorter than skeptics [4]. Thousands of people used the FutureSearch beta and ran >10k high-effort forecasts, on all sorts of diverse topics. Ask it anything about the future. We now support decision forecasts too: “If I do X, will I achieve this outcome?” Forecasting, as a capability, is useful even at the level of expert human crowds. But we predict that we will soon have strongly superhuman forecasting. People who bet against AI capability trend lines tend to lose, and the trend line in AI forecast accuracy tells a pretty clear story [5]. There’s no reason to think the best human forecasters have figured out everything predictable about the world. There’s a lot more signal to be found. And if you’re skeptical, try it. We’ve seen our fair share of exaggerated claims about AI forecasting accuracy [6]. So part of the reason we made the free tier give a few of our highest effort forecasters free is to let anyone verify the quality. [1] https://www.metaculus.com/tournament/summer-futureeval-2026/ [2] evals.futuresearch.ai [3] markets.futuresearch.ai [4] https://ai-2027.com/research/timelines-forecast [5] https://www.astralcodexten.com/p/the-ai-superforecasters-are... [6] https://www.lesswrong.com/posts/uGkRcHqatmPkvpGLq/contra-pap...

Revenue N/A
AI Tools
Vocab Top

Vocab Top

Show HN: Vocab Top – AI-powered vocabulary builder that helps you retain words

Revenue N/A
AI Tools
An AI agent that trades inside limits you set, starting on paper

An AI agent that trades inside limits you set, starting on paper

Show HN: An AI agent that trades inside limits you set, starting on paper

Revenue N/A
AI Tools
AuditBadger

AuditBadger

Hi, Wanted to share something I've been working on for over a year. AuditBadger is a compliance management platform that uses AI to write policies (there are underlying "templates" with basic requirements), rewrite controls (or trust service criterions) to match the company context, help figure out your own controls, does initial risk assessment, and business continuity planning (which at least gives you an example of how the process should look like). Fun fact - I wanted to share this a year ago, but then I spotted something similar here. The most common comment was about lacking the SOC 2 report, so I decided to pick the fight. I got SOC 2 Type I first, and then recently finished SOC 2 Type II using the tool alone. It took some time - both learning the process, the SOC 2 gotchas, and implementing automatic evidence collection. We're now adding support for the European AI Act and NIS 2; HIPAA is already there (though it requires me to explicitly enable it for customers who want to test it), and CyberEssentials and ENS are coming later this year. The platform is now complete, but my business partner (ISO 27001 Lead Auditor) and I are still dog-fooding it. Everything we build is either based on our own pain points or our customers'—most of them joined our Slack where we try to help them if they get stuck. If you have any questions, I'll be happy to answer them all.

Revenue N/A
AI Tools
ClickBench Playground

ClickBench Playground

I created it mostly for testing and exploration, but the main reason was that it became possible after previous work.

Revenue N/A
AI Tools
Spltty

Spltty

I built Spltty, a Ruby CLI for tracking shared expenses, custom splits, and settlements using plain Markdown files. The article explains how it evolved from a Claude-managed folder into a CLI

Revenue N/A
AI Tools
Silo

Silo

Show HN: Silo – S3-compatible object storage, a maintained fork of MinIO

Revenue N/A
AI Tools
Science for Kids

Science for Kids

There’s a really good kids’ science magazine called Oyla, but it’s aimed at kids 12+ I subscribed a couple of years ago thinking that it would be fine for my son (then 8yo, but an advanced reader) but, even now (10yo), the articles are too dense for him and don’t hold his interest. Last year, the same company launched ‘Oyla Junior’ aimed at younger kids ~8yo. But I read a couple of the articles and I think they’re too oversimplified. I wanted something in between, but didn’t find anything suitable online. So I used AI to create something, and did a bunch of iterations to tune the sentence length, the flow of the articles, how much info is introduced in each paragraph etc. I still want to make it better, but it’s already good enough for my son to read.

Revenue N/A
AI Tools
Hope, 1-on-1 AI Tutor for kids Grades 2-10 learning math

Hope, 1-on-1 AI Tutor for kids Grades 2-10 learning math

Hey HN! My friend and I have been working for the last 3 months on HOPE, a platform that provides students with 1 on 1 math support to improve their confidence and teach them by doing and asking questions. We would really appreciate anyone who's down to try it out with their kids and give their thoughts.

Revenue N/A