Independent AI Research · Mechanistic Interpretability · Sparse Autoencoders

NGNN-SAE: fully utilized feature dictionaries for LLM interpretability

NEGABO Research is developing NGNN-SAE, a proprietary closed-form-inspired feature learning approach for sparse autoencoders, focused on fully utilized feature dictionaries, atom-level statistics, and ablation-ready explanations.

We are currently looking for research partners with access to LLM activations to compare NGNN-SAE against established SAE baselines.

atom 001
atom 002
atom 003
atom 004
coarse layer
semantic layer
detail layer
full view
ablation Δ
usage entropy
channel signal
archetype
0

Dead features observed

Under current early-stage deadness criteria using our proprietary learning core.

atom

Level statistics

Utilization, reconstruction contribution, ablation impact, and channel diagnostics.

Δ

Ablation-ready

Designed to move from feature discovery toward causal explanation workflows.

coarse→detail

Rank-progressive analysis

Testing whether compact views capture coarse structure while larger views add detail.

The problem

Sparse autoencoders are promising — but reliability still matters.

Sparse autoencoders are increasingly used to decompose neural network activations into more interpretable features. But practical workflows still face difficult questions: Which features are wasted? Which features are stable? Which ones are causally useful? And how do we scale analysis without turning interpretability into another black box?

Dead or underused features

Capacity can be wasted by features that rarely or never participate.

Fragile explanations

A feature can look meaningful but fail under ablation, steering, or context shifts.

High evaluation cost

Large SAE suites require substantial activations, compute, and tooling.

Weak feature-to-mechanism bridge

Researchers need not only features, but stable statistics and causal tests.

Our approach

NGNN-SAE treats feature learning as a measurable, inspectable system.

Our current research combines a proprietary closed-form-inspired learning core, sparse and structured feature dictionaries, hierarchical representation analysis, and ablation-oriented evaluation. The goal is not only to reconstruct activations, but to produce features that are used, measured, compared, and tested.

Proprietary learning core

A closed-form-inspired procedure designed to stabilize feature utilization and produce inspectable feature dictionaries. Implementation details are withheld until IP and paper filings are complete.

Atom statistics

Tracking utilization, channel signals, reconstruction contribution, redundancy, and ablation impact.

Ablation-ready dictionaries

Features are evaluated by what changes when they are removed, not only by how plausible they look.

Rank-progressive analysis

Testing whether compact representations capture coarse structure while larger representations add semantic and instance-level detail.

Archetype graph dashboard

Codes can be organized into prototype families, core samples, boundary cases, typical atoms, and residual explanations.

Multibank stability

Independent feature banks can be aligned to distinguish stable atoms from training artifacts.

Current evidence

Early results suggest a useful research direction.

Now we want to validate it on real LLM activations.

The results below are early-stage and should be treated as hypotheses until independently benchmarked on large-scale LLM activation datasets.

Zero-dead-feature behavior under current criteria

Current tests indicate zero-dead-feature behavior under our defined criteria. We do not claim that utilization alone proves feature usefulness.

Useful atom statistics

We can measure utilization, redundancy, reconstruction contribution, channel stability, and ablation effect.

Rank-progressive structure

Early hierarchical experiments suggest that compact views can capture semantically useful structure, while larger views add detail.

Dashboard-ready XAI layer

Archetype graphs and multibank stability analysis provide a path from raw features to explainable, inspectable feature families.

Disclosure policy

Public evidence, private implementation.

Before patent and paper filings are complete, we publish the problem framing, benchmark design, evaluation metrics, and high-level results. We do not publish implementation details that would allow the core method to be reconstructed.

Public

Motivation, target use cases, benchmark tables, dead-feature criteria, utilization metrics, and high-level evidence.

Qualified partners

Deeper evaluation results, experimental setup, limitations, and collaboration-specific technical material under appropriate confidentiality terms.

After filings

Full technical disclosure can follow once the IP and publication process is ready for reproducible release.

Benchmark plan

Comparing NGNN-SAE against established SAE baselines.

We are especially interested in benchmarks where NGNN-SAE can be evaluated not just by reconstruction loss, but by feature utilization, ablation usefulness, semantic stability, and transfer across layers, datasets, and model families.

MetricStandard SAETopK SAEJumpReLU SAEFeature Choice SAENGNN-SAE
Dead feature rateTBDTBDTBDTBDTBD
Near-dead feature rateTBDTBDTBDTBDTBD
Reconstruction lossTBDTBDTBDTBDTBD
L0 / average active featuresTBDTBDTBDTBDTBD
Feature utilization entropyTBDTBDTBDTBDTBD
Functional ablation impactTBDTBDTBDTBDTBD
Training time / computeTBDTBDTBDTBDTBD
Interpretability examplesTBDTBDTBDTBDTBD
Research partners

Looking for partners with LLM activations, interpretability questions, or SAE baselines.

We are a small research startup. The best next milestone is a clean benchmark collaboration with one model, one layer, and clearly agreed metrics.

Select scope

Choose one model, one layer, and one activation dataset.

Agree metrics

Define baselines, deadness criteria, utilization, ablation, and compute limits.

Run comparison

Train NGNN-SAE and baseline SAEs under comparable conditions.

Analyze features

Inspect atom statistics, ablations, archetypes, and stability.

Report results

Create a private report, public note, or co-authored technical brief.

Research notes

Initial publication queue.

Short, falsifiable research notes are better than broad hype posts.

Why dead features matter in sparse autoencoders

A practical framing of wasted capacity and reliability.

From zero dead features to useful features

Why utilization is necessary but not sufficient.

Atom statistics beyond reconstruction loss

Tracking utilization, redundancy, and ablation effect.

Coarse-to-fine representation analysis

How compact views and larger views differ in reconstruction, neighborhoods, and explanations.

Archetype Graphs

Turning feature codes into interpretable prototype families.

Call for partners

Benchmarking NGNN-SAE on LLM activations.

Have access to LLM activations?

We are looking for partners to run a clean comparison between NGNN-SAE and current SAE baselines. The goal is simple: determine whether NGNN-SAE produces more fully utilized, measurable, and ablation-useful feature dictionaries without exposing proprietary implementation details before filings are complete.

Contact NEGABO Research