Modules
The main public API is exposed from the djinn package:
from djinn import DJINN_Regressor, DJINN_Classifier
In particular, djinn.DJINN_Regressor
is where the magic happens.
djinn
DJINN public package interface.
This package exports the high-level djinn.djinn module, which contains
the public regression/classification APIs and model persistence helpers.
Modules
- djinn
Public model classes and loading helpers.
- class djinn.DJINN_Classifier(n_trees=1, max_tree_depth=4, dropout_keep_prob=1.0, **kwargs)[source]
Bases:
DJINN_RegressorDJINN classification model.
Inherits all training, saving, and loading behaviour from
DJINN_Regressor. The only behavioural difference is inbayesian_predict(), where no output scaling is applied andnp.argmaxis used to convert softmax distributions into class predictions.- Parameters:
n_trees (int, optional) – Number of trees in the random forest (equal to the number of neural networks).
max_tree_depth (int, optional) – Maximum depth of decision tree. The neural network has
max_tree_depth - 1hidden layers.dropout_keep_prob (float, optional) – Probability of keeping a neuron in dropout layers.
**kwargs – Optional keyword arguments forwarded to
DJINN_Regressor.
- bayesian_predict(x_test, n_iters, seed=None)[source]
Bayesian distribution of class predictions for a set of test inputs.
Evaluates each tree network
n_iterstimes (with dropout active) to build a predictive distribution over class probabilities, then returns theargmaxof the 25th, 50th, and 75th percentiles as integer class labels alongside the raw sample dictionary.- Parameters:
- Returns:
If
n_itersisNone, returns a 1-D array of predicted class indices with shape(n_test,). Otherwise returns(lower, middle, upper, samples), where percentile outputs are 1-D arrays of class indices andsamplescontains per-tree probability draws.- Return type:
ndarray or tuple
- predict(x_test, seed=None)[source]
Predict class labels for a set of test inputs.
Calls
bayesian_predict()withn_iters=None(single deterministic forward pass per network) and returns theargmaxclass predictions.- Parameters:
x_test (ndarray) – Input feature matrix for testing.
seed (int or None, optional) – Random seed for reproducibility.
- Returns:
Predicted class index for each test point, shape
(n_test,).- Return type:
ndarray
- class djinn.DJINN_Regressor(n_trees=1, max_tree_depth=4, dropout_keep_prob=1.0, **kwargs)[source]
Bases:
objectDJINN regression model (PyTorch backend).
- Parameters:
n_trees (int, optional) – Number of trees in the random forest (equal to the number of neural networks).
max_tree_depth (int, optional) – Maximum depth of decision tree. The neural network has
max_tree_depth - 1hidden layers.dropout_keep_prob (float, optional) – Probability of keeping a neuron in dropout layers.
**kwargs – Optional preloaded state including scalers, models, paths, and device.
- bayesian_predict(x_test, n_iters, seed=None)[source]
Bayesian distribution of predictions for a set of test inputs.
Evaluates each tree network
n_iterstimes (with dropout active) to build a predictive distribution, then returns the 25th, 50th, and 75th percentiles alongside the raw sample dictionary.- Parameters:
- Returns:
If
n_itersisNone, returns mean predictions with shape(n_test, n_outputs). Otherwise returns(lower, middle, upper, samples), where percentile arrays have shape(n_test, n_outputs)andsamplescontains per-tree prediction draws.- Return type:
ndarray or tuple
- bma_predict(x_test, n_iters=100, seed=None)[source]
Return Bayesian model averaging samples and summary statistics.
- Parameters:
- Returns:
Dictionary containing percentile summaries and stacked prediction samples under
predictionswith shape(n_iters * n_trees, n_test, n_outputs).- Return type:
- collect_tree_predictions(predictions)[source]
Gather and reshape the full distribution of per-tree predictions.
- Parameters:
predictions (dict) –
"predictions"sub-dictionary from the dictionary returned bybayesian_predict().- Returns:
Reshaped predictions with shape
(n_iters * n_trees, n_test, n_outputs).- Return type:
ndarray
- continue_training(X, Y, training_epochs, learning_rate, batch_size, learn_rate=None, seed=None)[source]
Continue training an existing model (must call
load_model()first).Delegates to
neural_network.torch_continue_trainingand re-saves each tree checkpoint in place.- Parameters:
X (ndarray) – Input feature matrix for training.
Y (ndarray) – Target array for training.
training_epochs (int) – Additional epochs to train.
learning_rate (float) – Learning rate.
learn_rate (float or None, optional) – Backward-compatible alias for
learning_rate.batch_size (int) – Number of samples per batch.
seed (int or None, optional) – Random seed for reproducibility.
- Return type:
None
- fit(X, Y, epochs=None, learning_rate=None, learn_rate=None, batch_size=None, weight_decay=1e-08, save_files=True, save_model=True, model_name='djinn_model', model_path='./', seed=None)[source]
Train DJINN, auto-selecting hyperparameters when not supplied.
If
learning_rateis None, callsget_hyperparameters()first and uses the returned values before delegating totrain().- Parameters:
X (ndarray) – Input feature matrix for training.
Y (ndarray) – Target array for training.
epochs (int or None, optional) – Number of training epochs.
learning_rate (float or None, optional) – Learning rate for weight and bias optimization. If
None, hyperparameters are tuned automatically.learn_rate (float or None, optional) – Backward-compatible alias for
learning_rate.batch_size (int or None, optional) – Number of samples per batch.
weight_decay (float, optional) – Multiplier for L2 penalty on weights.
save_files (bool, optional) – If
True, saves train/validation cost and weights.save_model (bool, optional) – If
True, saves the trained model.model_name (str, optional) – File name for the model.
model_path (str, optional) – Directory where model/files are saved.
seed (int or None, optional) – Random seed for reproducibility.
- Return type:
None
- classmethod from_json(json_path)[source]
Reconstruct a DJINN_Regressor from a saved JSON state file.
Restores all hyperparameters and scalers so the instance is ready for
load_model(),predict(), orcontinue_training().- Parameters:
json_path (str or pathlib.Path) – Path to the
.jsonfile written bytrain().- Returns:
Restored regressor instance.
- Return type:
- get_hyperparameters(X, Y, weight_decay=1e-08, seed=None)[source]
Automatically select DJINN hyperparameters.
Returns learning rate, number of epochs, and batch size by running a short auto-tuning search using the PyTorch training utilities in
neural_network.py.- Parameters:
- Raises:
Exception – If a decision tree cannot be built from the data.
- Returns:
Dictionary with keys
batch_size,learning_rate, andepochs.- Return type:
- load_model(model_name, model_path)[source]
Reload PyTorch checkpoints for a saved model.
Restores each tree’s
.ptcheckpoint from disk usingneural_network.load_tree_model.- Parameters:
model_name (str) – Name of the saved model directory.
model_path (str or pathlib.Path) – Parent directory that contains the model folder.
- Return type:
None
- predict(x_test, seed=None)[source]
Predict target values for a set of test inputs.
Calls
bayesian_predict()withn_iters=None(single deterministic forward pass per network) and returns the mean.- Parameters:
x_test (ndarray) – Input feature matrix for testing.
seed (int or None, optional) – Random seed for reproducibility.
- Returns:
Mean target value prediction for each test point, shape
(n_test, n_outputs).- Return type:
ndarray
- save(model_path, overwrite=False)[source]
Persist the currently loaded model under an explicit output path.
Checkpoints are written from the in-memory models, so this works whether or not
train()was called withsave_model=True.- Parameters:
model_path (str or pathlib.Path) – Target base path. Writes checkpoints to
<model_path>/and metadata to<model_path>.json.overwrite (bool, optional) – If
True, delete and replacemodel_pathwhen it already exists. Defaults toFalse, which raises instead of silently deleting an existing directory.
- Returns:
Saved model directory path.
- Return type:
- Raises:
RuntimeError – If there are no trained or loaded models to save.
FileExistsError – If
model_pathalready exists andoverwriteisFalse.
- train(X, Y, epochs=1000, learning_rate=0.001, learn_rate=None, batch_size=0, weight_decay=1e-08, save_files=True, save_model=True, model_name='djinn_model', model_path='./', ntrees=None, seed=None, eval_every=1)[source]
Train DJINN with specified hyperparameters.
Builds a random forest, maps each tree to a PyTorch MLP via
random_forest.tree_to_nn_weights, then trains every network usingneural_network.torch_dropout_regression.- Parameters:
X (ndarray) – Input feature matrix for training.
Y (ndarray) – Target array for training.
epochs (int, optional) – Number of training epochs.
learning_rate (float, optional) – Learning rate for weight and bias optimization.
learn_rate (float or None, optional) – Backward-compatible alias for
learning_rate.batch_size (int, optional) – Number of samples per batch. If
0, uses 5% of the dataset.weight_decay (float, optional) – Multiplier for L2 penalty on weights.
save_files (bool, optional) – If
True, saves train/validation cost per epoch and weights/biases.save_model (bool, optional) – If
True, saves the trained model.model_name (str, optional) – File name for the model when
save_modelisTrue.model_path (str, optional) – Directory where model/files are saved.
seed (int or None, optional) – Random seed for reproducibility.
- Raises:
Exception – If a decision tree cannot be built from the data.
- Return type:
None
- djinn.load(model_path)[source]
Load a saved DJINN model from path.
- Parameters:
model_path (str or pathlib.Path) – Path to the model directory or its JSON sidecar.
- Returns:
Reconstructed model with checkpoints loaded.
- Return type: