Against the field#

How the models here compare with what people actually use on tabular data: XGBoost, LightGBM, Random Forest, a two-layer MLP, and two differentiable tree ensembles from the literature, GRANDE (Marton et al., ICLR 2024) and NODE (Popov et al., ICLR 2020). Twenty-four datasets: Iris, Wine, Breast Cancer, Digits and twenty from the OpenML CC-18 suite, capped at 5 000 rows by stratified subsampling. Three seeds of stratified five-fold cross-validation, features standardised on the training fold, accuracy averaged over the fifteen folds. Every model in its own process, one BLAS thread.

Nothing is tuned. Each model runs one fixed configuration on every dataset: XGBoost 300 trees of depth 6 at learning rate 0.1, LightGBM 300 trees of 31 leaves, Random Forest 300 unbounded trees, MLP 64-64 for 300 iterations, GRANDE 256 trees of depth 5 with its published defaults, NODE one layer of 256 oblivious trees of depth 6 for 50 epochs with early stopping, the soft tree at depth 4 for 150 epochs, the per-leaf tree at depth 6 for 180 epochs, GAL for 150 epochs with residual initialisation. A tuned XGBoost would do better on the small datasets, where 300 deep trees overfit 150 rows; so would a tuned soft tree. The comparison is between defaults, which is what a first run looks like, and it should be read that way.

The script is benchmarks/rakipler.py; benchmarks/rakipler_ozet.py turns its output into the tables below. Bold marks the best in a row.

Dataset

n

K

SoftTree d4

SoftTree per-leaf

GAL

RandomForest

MLP

XGBoost

LightGBM

GRANDE

NODE

Iris

150

3

0.958

0.960

0.956

0.949

0.951

0.940

0.953

0.956

0.873

Wine

178

3

0.978

0.972

0.985

0.978

0.981

0.948

0.972

0.983

0.729

Cancer

569

2

0.970

0.974

0.975

0.958

0.975

0.965

0.966

0.970

0.828

Digits

1797

10

0.925

0.927

0.928

0.977

0.974

0.966

0.974

0.964

0.879

banknote-authentication

1372

2

1.000

0.989

0.989

0.993

1.000

0.996

0.996

0.999

0.934

blood-transfusion-service-center

748

2

0.799

0.785

0.799

0.740

0.794

0.742

0.737

0.787

0.762

phoneme

5000

2

0.849

0.836

0.847

0.908

0.879

0.895

0.900

0.881

0.825

diabetes

768

2

0.731

0.767

0.762

0.761

0.742

0.730

0.726

0.754

0.683

vehicle

846

4

0.798

0.780

0.803

0.749

0.838

0.759

0.773

0.761

0.465

segment

2310

7

0.948

0.952

0.942

0.971

0.965

0.975

0.979

0.965

0.863

satimage

5000

6

0.880

0.878

0.892

0.915

0.907

0.916

0.923

0.895

0.835

optdigits

5000

10

0.937

0.938

0.932

0.982

0.981

0.976

0.981

0.975

0.943

pendigits

5000

10

0.930

0.954

0.960

0.988

0.993

0.986

0.988

0.984

0.921

spambase

4601

2

0.929

0.930

0.929

0.954

0.939

0.955

0.958

0.951

0.924

steel-plates-fault

1941

2

1.000

1.000

0.999

0.994

0.999

1.000

1.000

1.000

0.994

ilpd

583

2

0.705

0.716

0.721

0.707

0.714

0.694

0.699

0.722

0.714

climate-model-simulation-crashes

540

2

0.932

0.944

0.959

0.920

0.957

0.950

0.944

0.948

0.915

texture

5000

11

0.938

0.993

0.984

0.976

0.998

0.981

0.986

0.974

0.840

wall-robot-navigation

5000

4

0.875

0.876

0.903

0.993

0.922

0.996

0.996

0.987

0.826

qsar-biodeg

1055

2

0.855

0.872

0.870

0.871

0.872

0.867

0.862

0.868

0.682

ozone-level-8hr

2534

2

0.929

0.941

0.940

0.944

0.938

0.944

0.945

0.942

0.937

kc1

2109

2

0.853

0.854

0.859

0.862

0.857

0.858

0.856

0.856

0.846

pc1

1109

2

0.921

0.931

0.931

0.939

0.930

0.934

0.935

0.935

0.931

pc3

1563

2

0.875

0.897

0.894

0.899

0.879

0.892

0.891

0.888

0.898

mean accuracy

0.896

0.903

0.907

0.914

0.916

0.911

0.914

0.914

0.835

mean rank

6.27

5.06

4.52

4.12

3.79

4.66

4.16

4.29

8.12

mean fit time (s)

4.4

18.5

1.6

0.6

32.6

0.3

0.3

110.3

121.6

Paired against XGBoost#

Difference in accuracy points, model minus XGBoost, over the datasets; a tie is a difference within half a point; the Wilcoxon signed-rank test is over the per-dataset means.

Model

mean accuracy

mean rank

vs XGBoost (points)

win/tie/loss

Wilcoxon p

SoftTree d4

0.896

6.27

-1.46

5/5/14

0.049

SoftTree per-leaf

0.903

5.06

-0.83

10/4/10

0.422

GAL

0.907

4.52

-0.44

8/7/9

0.726

RandomForest

0.914

4.12

+0.26

9/11/4

0.317

MLP

0.916

3.79

+0.51

11/6/7

0.264

XGBoost

0.911

4.66

+0.00

0/24/0

LightGBM

0.914

4.16

+0.32

8/13/3

0.030

GRANDE

0.914

4.29

+0.33

6/13/5

1.000

NODE

0.835

8.12

-7.58

3/1/20

0.000

Over all twenty-four datasets the gradient-boosted ensembles, Random Forest and the MLP are the top group and the soft models sit a point or two below them; NODE, with these defaults and this budget, is last everywhere. That is the expected result and it is not the interesting one.

Small datasets#

Split by size, the picture reverses. On the datasets with at most 1 000 rows:

Model

mean accuracy

mean rank

vs XGBoost (points)

win/tie/loss

Wilcoxon p

SoftTree d4

0.859

4.44

+1.77

5/2/1

0.078

SoftTree per-leaf

0.862

3.62

+2.14

7/0/1

0.016

GAL

0.870

1.81

+2.91

8/0/0

0.008

RandomForest

0.845

6.56

+0.43

4/1/3

0.547

MLP

0.869

3.25

+2.81

8/0/0

0.008

XGBoost

0.841

7.00

+0.00

0/8/0

LightGBM

0.846

6.62

+0.53

3/3/2

0.383

GRANDE

0.860

3.62

+1.93

6/2/0

0.016

NODE

0.746

8.06

-9.48

2/0/6

0.039

and on the rest:

Model

mean accuracy

mean rank

vs XGBoost (points)

win/tie/loss

Wilcoxon p

SoftTree d4

0.915

7.18

-3.08

0/3/13

0.000

SoftTree per-leaf

0.923

5.78

-2.31

3/4/9

0.009

GAL

0.925

5.88

-2.11

0/7/9

0.003

RandomForest

0.948

2.91

+0.18

5/10/1

0.348

MLP

0.940

4.06

-0.64

3/6/7

0.298

XGBoost

0.946

3.49

+0.00

0/16/0

LightGBM

0.948

2.93

+0.22

5/10/1

0.023

GRANDE

0.941

4.62

-0.46

0/11/5

0.005

NODE

0.880

8.15

-6.62

1/1/14

0.000

On the small datasets every gradient-trained smooth model beats XGBoost: GAL on every one of them, the per-leaf soft tree on seven of eight, the MLP on every one. Boosting with 300 deep trees has more capacity than a few hundred rows can constrain, and a smooth model with a few dozen parameters per node does not. Above 1 000 rows, sixteen datasets, the order flips and XGBoost, LightGBM and Random Forest win almost every comparison against the soft models.

Two cautions. The split at 1 000 rows was chosen after seeing the data, and eight datasets is a small sample, so the p-values in the small-data table are exploratory, not confirmatory. And the MLP wins there too, so the finding is about smooth gradient-trained models against untuned boosting on small tables, not about trees in particular. What the soft tree adds over the MLP in that regime is what it adds everywhere: a model that can be read (Explaining a prediction), exported as rules or as plain numpy, and whose probabilities are calibrated (User guide).

Averaging soft trees#

The natural objection to the large-data result is that it compares one soft tree with 300 boosted ones. benchmarks/soft_forest.py answers it: 25 per-leaf soft trees, each on a bootstrap sample with 90 epochs, probabilities averaged, on the sixteen datasets above 1 000 rows under the same folds and seeds. Nothing tuned here either.

dataset

n

K

soft forest (25 trees)

per-leaf soft tree

Random Forest

XGBoost

LightGBM

forest s/fit

banknote-authentication

1372

2

0.987 ± 0.007

0.989

0.993

0.996

0.996

21

kc1

2109

2

0.858 ± 0.011

0.854

0.862

0.858

0.856

10

ozone-level-8hr

2534

2

0.946 ± 0.006

0.941

0.944

0.944

0.945

12

pc1

1109

2

0.933 ± 0.007

0.931

0.939

0.934

0.935

10

pc3

1563

2

0.898 ± 0.007

0.897

0.899

0.892

0.891

9

phoneme

5000

2

0.856 ± 0.015

0.836

0.908

0.895

0.900

56

qsar-biodeg

1055

2

0.883 ± 0.015

0.872

0.871

0.867

0.862

8

spambase

4601

2

0.937 ± 0.005

0.930

0.954

0.955

0.958

109

steel-plates-fault

1941

2

1.000 ± 0.001

1.000

0.994

1.000

1.000

61

wall-robot-navigation

5000

4

0.908 ± 0.012

0.876

0.993

0.996

0.996

66

satimage

5000

6

0.890 ± 0.009

0.878

0.915

0.916

0.923

46

segment

2310

7

0.952 ± 0.009

0.952

0.971

0.975

0.979

34

Digits

1797

10

0.976 ± 0.007

0.927

0.977

0.966

0.974

18

optdigits

5000

10

0.978 ± 0.005

0.938

0.982

0.976

0.981

47

pendigits

5000

10

0.985 ± 0.004

0.954

0.988

0.986

0.988

68

texture

5000

11

0.995 ± 0.003

0.993

0.976

0.981

0.986

474

Averaging closes about half of the gap and no more. Against XGBoost the single per-leaf tree is 2.3 points behind on average (median 1.5, wins, ties, losses 3/4/9, Wilcoxon p = 0.011); the forest is 1.0 point behind (median 0.05, 4/6/6, p = 0.32), which the test can no longer distinguish from XGBoost, while Random Forest and LightGBM sit 0.2 points ahead of it. The forest beats the single tree on nine datasets and loses on none (p = 0.002). The gain is concentrated where the tree was weakest: on the seven multi-class datasets the gap shrinks from 4.0 points to 1.6, on the nine binary ones from 1.0 to 0.5. Where boosting wins clearly it keeps winning: wall-robot-navigation stays 8.8 points behind, phoneme 3.9.

The price is the reason to use a soft tree in the first place. A forest has no single path to read, no rule list and no per-prediction counterfactual; in that respect it is a Random Forest that trains 160 times slower (65 s per fit against 0.4 s for XGBoost and 27 s for the single tree) for the same accuracy. It is reported here so that the question has a measured answer, not as a recommendation: if the explanation does not matter, use LightGBM; if it does, use one tree and accept the two points.

So the honest summary for someone choosing a model: with a few hundred rows and a need to explain the predictions, a soft tree or GAL is a reasonable first choice and will likely match or beat an untuned boosting model. With thousands of rows and accuracy as the only criterion, use LightGBM.