NGNN-SAE: fully utilized feature dictionaries for LLM interpretability
NEGABO Research is developing NGNN-SAE, a proprietary closed-form-inspired feature learning approach for sparse autoencoders, focused on fully utilized feature dictionaries, atom-level statistics, and ablation-ready explanations.
We are currently looking for research partners with access to LLM activations to compare NGNN-SAE against established SAE baselines.
Dead features observed
Under current early-stage deadness criteria using our proprietary learning core.
Level statistics
Utilization, reconstruction contribution, ablation impact, and channel diagnostics.
Ablation-ready
Designed to move from feature discovery toward causal explanation workflows.
Rank-progressive analysis
Testing whether compact views capture coarse structure while larger views add detail.
Sparse autoencoders are promising — but reliability still matters.
Sparse autoencoders are increasingly used to decompose neural network activations into more interpretable features. But practical workflows still face difficult questions: Which features are wasted? Which features are stable? Which ones are causally useful? And how do we scale analysis without turning interpretability into another black box?
Dead or underused features
Capacity can be wasted by features that rarely or never participate.
Fragile explanations
A feature can look meaningful but fail under ablation, steering, or context shifts.
High evaluation cost
Large SAE suites require substantial activations, compute, and tooling.
Weak feature-to-mechanism bridge
Researchers need not only features, but stable statistics and causal tests.
NGNN-SAE treats feature learning as a measurable, inspectable system.
Our current research combines a proprietary closed-form-inspired learning core, sparse and structured feature dictionaries, hierarchical representation analysis, and ablation-oriented evaluation. The goal is not only to reconstruct activations, but to produce features that are used, measured, compared, and tested.
Proprietary learning core
A closed-form-inspired procedure designed to stabilize feature utilization and produce inspectable feature dictionaries. Implementation details are withheld until IP and paper filings are complete.
Atom statistics
Tracking utilization, channel signals, reconstruction contribution, redundancy, and ablation impact.
Ablation-ready dictionaries
Features are evaluated by what changes when they are removed, not only by how plausible they look.
Rank-progressive analysis
Testing whether compact representations capture coarse structure while larger representations add semantic and instance-level detail.
Archetype graph dashboard
Codes can be organized into prototype families, core samples, boundary cases, typical atoms, and residual explanations.
Multibank stability
Independent feature banks can be aligned to distinguish stable atoms from training artifacts.
Early results suggest a useful research direction.
Now we want to validate it on real LLM activations.
Zero-dead-feature behavior under current criteria
Current tests indicate zero-dead-feature behavior under our defined criteria. We do not claim that utilization alone proves feature usefulness.
Useful atom statistics
We can measure utilization, redundancy, reconstruction contribution, channel stability, and ablation effect.
Rank-progressive structure
Early hierarchical experiments suggest that compact views can capture semantically useful structure, while larger views add detail.
Dashboard-ready XAI layer
Archetype graphs and multibank stability analysis provide a path from raw features to explainable, inspectable feature families.
Public evidence, private implementation.
Before patent and paper filings are complete, we publish the problem framing, benchmark design, evaluation metrics, and high-level results. We do not publish implementation details that would allow the core method to be reconstructed.
Public
Motivation, target use cases, benchmark tables, dead-feature criteria, utilization metrics, and high-level evidence.
Qualified partners
Deeper evaluation results, experimental setup, limitations, and collaboration-specific technical material under appropriate confidentiality terms.
After filings
Full technical disclosure can follow once the IP and publication process is ready for reproducible release.
Comparing NGNN-SAE against established SAE baselines.
We are especially interested in benchmarks where NGNN-SAE can be evaluated not just by reconstruction loss, but by feature utilization, ablation usefulness, semantic stability, and transfer across layers, datasets, and model families.
| Metric | Standard SAE | TopK SAE | JumpReLU SAE | Feature Choice SAE | NGNN-SAE |
|---|---|---|---|---|---|
| Dead feature rate | TBD | TBD | TBD | TBD | TBD |
| Near-dead feature rate | TBD | TBD | TBD | TBD | TBD |
| Reconstruction loss | TBD | TBD | TBD | TBD | TBD |
| L0 / average active features | TBD | TBD | TBD | TBD | TBD |
| Feature utilization entropy | TBD | TBD | TBD | TBD | TBD |
| Functional ablation impact | TBD | TBD | TBD | TBD | TBD |
| Training time / compute | TBD | TBD | TBD | TBD | TBD |
| Interpretability examples | TBD | TBD | TBD | TBD | TBD |
Looking for partners with LLM activations, interpretability questions, or SAE baselines.
We are a small research startup. The best next milestone is a clean benchmark collaboration with one model, one layer, and clearly agreed metrics.
Select scope
Choose one model, one layer, and one activation dataset.
Agree metrics
Define baselines, deadness criteria, utilization, ablation, and compute limits.
Run comparison
Train NGNN-SAE and baseline SAEs under comparable conditions.
Analyze features
Inspect atom statistics, ablations, archetypes, and stability.
Report results
Create a private report, public note, or co-authored technical brief.
Initial publication queue.
Short, falsifiable research notes are better than broad hype posts.
Why dead features matter in sparse autoencoders
A practical framing of wasted capacity and reliability.
From zero dead features to useful features
Why utilization is necessary but not sufficient.
Atom statistics beyond reconstruction loss
Tracking utilization, redundancy, and ablation effect.
Coarse-to-fine representation analysis
How compact views and larger views differ in reconstruction, neighborhoods, and explanations.
Archetype Graphs
Turning feature codes into interpretable prototype families.
Call for partners
Benchmarking NGNN-SAE on LLM activations.
Have access to LLM activations?
We are looking for partners to run a clean comparison between NGNN-SAE and current SAE baselines. The goal is simple: determine whether NGNN-SAE produces more fully utilized, measurable, and ablation-useful feature dictionaries without exposing proprietary implementation details before filings are complete.