| Name |
Format |
Description |
Link |
|
21 |
The data contained herein are five input features (i.e., heat flow, distance to the nearest quaternary fault, distance to the nearest quaternary magma body, seismic event density, maximum horizontal stress) and labels (i.e., where known geothermal systems have been identified) from Williams and DeAngelo (2008) and nine favorability maps from Mordensky et al. (2023). The favorability maps are the untransformed predictions from models resulting from the features and labels used with either the methods presented in Williams and DeAngelo (2008) or the machine learning approaches presented in Mordensky et al. (2023). |
https://doi.org/10.5066/P9V1Q9XM |
|
0 |
Our study aims to develop robust data-driven methods with the goals of reducing bias and improving predictive ability. We present and compare nine favorability maps for geothermal resources in the western United States using data from the U.S. Geological Survey's 2008 geothermal resource assessment. Two favorability maps are created using the expert decision-dependent methods from the 2008 assessment (i.e., weight-of-evidence and logistic regression). With the same data, we then create six different favorability maps using logistic regression (without underlying expert decisions), XGBoost, and support-vector machines paired with two training strategies. The training strategies are customized to address the inherent challenges of applying machine learning to the geothermal training data, which have no negative examples and severe class imbalance. We also create another favorability map using an artificial neural network. |
https://doi.org/10.1016/j.geothermics.2023.102662 |
|
21 |
Our research group of geoscientists and machine learning experts presents a process to help geoscientists understand the fundamentals of supervised learning by describing the general workflow (i.e., a conceptual pipeline) for supervised learning that must be understood by all the parties involved in a geoscience-machine learning endeavor. Terms critical for machine learning are introduced, defined, and used within the context of an overly simplified mock hydrological study to illustrate their appropriate usage, and then used again in the context of a published geothermal-machine learning study. |
https://www.geothermal-library.org/index.php?mode=pubs&action=view&record=1034680 |
|
21 |
This study aims to reduce expert input through robust data-driven analyses and better-suited data science techniques, with the goals of saving time, reducing bias, and improving predictive ability. We present six favorability maps for geothermal resources in the western United States created using two strategies applied to three modern machine learning algorithms (logistic regression, support-vector machines, and XGBoost). To provide a direct comparison to previous assessments, we use the same input data as the 2008 U.S. Geological Survey (USGS) conventional moderate- to high-temperature geothermal resource assessment. |
https://pangea.stanford.edu/ERE/db/IGAstandard/record_detail.php?id=35430 |
|
21 |
Recent evaluation of strategies for conventional hydrothermal resource assessment in the United States has relied upon machine learning methods (i.e., logistic regression, SVMs, XGBoost, and multilayer perceptron neural networks [i.e., MLPs]) to predict resource favorability using features (i.e., heat flow, distance to faults, distance to magma bodies, maximum horizontal strain, and seismic event density) from the U.S. Geological Survey’s 2008 Geothermal Resource Assessment. Two of the machine learning algorithms (i.e., SVMs and MLPs) must rely on model-agnostic measures of feature importance (i.e., measures of feature importance that are applicable regardless of an algorithm’s conceptual framework; e.g., sensitivity analyses and SHapely Additive exPlanation [SHAP] values), while the other two machine learning algorithms also offer straightforward, model-gnostic (i.e., algorithm-specific) measures to interpret the relative contributions of features on favorability predictions (i.e., feature coefficients for logistic regression, weight, gain, cover, and F score for XGBoost). Relative feature importance is measured for all machine learning algorithms using the model-agnostic measures, and, when possible, model-gnostic measures are shown for comparison. |
https://doi.org/10.1130/abs/2022AM-376621 |
|
21 |
To facilitate comparison of methods, we use the same data from the 2008 geothermal resource assessment (e.g., heat flow, horizontal stress) to train models from modern machine learning algorithms (i.e., logistic regression, eXtreme Gradient Boosting, support vector machines, and multilayer perceptron neural networks), which minimize dependence upon expert decisions. While some algorithms are simple (e.g., logistic regression), other algorithms are highly sophisticated (e.g., the neural network). Despite the contrast in complexity, the results from the very simple and highly complex algorithms are similar. |
https://doi.org/10.1130/abs/2022AM-377146 |
|
21 |
This study demonstrates that two foundational machine learning algorithms (logistic regression and XGBoost), implemented using unbiased data analysis strategies, agree with previous studies that relied much more heavily on expert-systems knowledge. |
https://doi.org/10.1130/abs/2021AM-365177 |