# Fileset

[Advanced Intelligent Systems - 2024 - Sato - Target Material Property‐Dependent Cluster Analysis of Inorganic Compounds.pdf](https://mdr.nims.go.jp/filesets/7a1366cc-60c5-48b3-bedd-6450cf6a0b40/download)

## Creator

[Nobuya Sato](https://orcid.org/0000-0002-8661-0410), [Akira Takahashi](https://orcid.org/0000-0002-3159-9007), [Shin Kiyohara](https://orcid.org/0000-0003-2890-5760), [Kei Terayama](https://orcid.org/0000-0003-3914-248X), [Ryo Tamura](https://orcid.org/0000-0002-0349-358X), [Fumiyasu Oba](https://orcid.org/0000-0001-7178-5333)

## Rights

[Creative Commons BY Attribution 4.0 International](https://creativecommons.org/licenses/by/4.0/)

## Other metadata

[Target Material Property‐Dependent Cluster Analysis of Inorganic Compounds](https://mdr.nims.go.jp/datasets/4022f7ed-bcdb-4f3b-b544-7feae8bcf6b1)

## Fulltext

Target Material Property‐Dependent Cluster Analysis of Inorganic CompoundsTarget Material Property-Dependent Cluster Analysis ofInorganic CompoundsNobuya Sato,* Akira Takahashi,* Shin Kiyohara, Kei Terayama, Ryo Tamura,and Fumiyasu Oba1. IntroductionIn the field of materials science, it is quite common to classifysubstances into various categories based on their constituent ele-ments and crystal structure characteristics. For instance, weoften categorize substances into oxides,sulfides, nitrides, or others according totheir composition and constituent ele-ments, or classify them into other classessuch as II–VI and III–V semiconductorsemploying other criteria. Furthermore, itis also a prevalent practice to classify mate-rials into prototype crystal structures suchas rock-salt or perovskite types. The funda-mental characteristics of constituent ele-ments and crystal structures that definematerials are extremely diverse, and theappropriate classification varies dependingon the intended application of the materi-als. It is essential to identify promising clas-ses of materials by considering materialproperties and functions closely relevantto the target application.Meanwhile, progress in machine learn-ing has brought new approaches to materi-als science.[1–4] One of the most commonapplications of machine learning wouldbe to predict material properties of interestfrom basic characteristics usually derived from constituentelements and/or crystal structures, which are often calleddescriptors. Interpretable or explainable machine learningtechniques[5–7] can unveil chemical trends and identify control-ling factors of material properties. For instance, the importanceN. Sato, A. Takahashi, S. Kiyohara, F. ObaLaboratory for Materials and StructuresInstitute of Innovative ResearchTokyo Institute of TechnologyR3-7, 4259 Nagatsuta, Midori-ku, 226-8501, JapanE-mail: nobuya.sato.000@gmail.com; takahashi.a.bb@m.titech.ac.jpS. KiyoharaInstitute for Materials ResearchTohoku University2-2-1 Katahira, Aoba-ku, Sendai 980-8577, JapanThe ORCID identification number(s) for the author(s) of this articlecan be found under https://doi.org/10.1002/aisy.202400253.© 2024 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbH. This is an open access article under the terms of the CreativeCommons Attribution License, which permits use, distribution andreproduction in any medium, provided the original work is properly cited.DOI: 10.1002/aisy.202400253K. TerayamaGraduate School of Medical Life ScienceYokohama City University1-7-29 Suehiro-cho, Tsurumi-ku, 230-0045, JapanK. TerayamaRIKEN Center for Advanced Intelligence Project1-4-1 Nihonbashi, Chuo-ku, Tokyo 103-0027, JapanK. Terayama, F. ObaMDX Research Center for Element StrategyInternational Research Frontiers InitiativeTokyo Institute of TechnologySE-6, 4259 Nagatsuta, Midori-ku, Yokohama 226-8501, JapanR. TamuraCenter for Basic Research on MaterialsNational Institute for Materials Science1-1 Namiki, Tsukuba 305-0044, JapanR. TamuraGraduate School of Frontier SciencesThe University of Tokyo5-1-5 Kashiwa-no-ha, Kashiwa 277-8568, JapanThe cluster analysis of materials categorizes them according to similarities basedon the features of materials, providing insight into the relationship between thematerials. Conventional cluster analyses typically use basic features derived fromthe chemical composition and crystal structure without considering target materialproperties such as the bandgap and dielectric constant. However, such approachesdo not meet demands for grading materials according to properties of interestsimultaneously with chemical and structural similarities. Herein, a clusteringmethod grouping similar materials in terms of both the target properties and basicfeatures is proposed. The clustering is compared considering the cohesive energywith that considering the bandgap of metal oxides, showing that their categori-zations are clearly different. Further, several clusters classified by the bandgap areanalyzed, and coordination environments related to each range of the bandgap arerevealed. The clustering for the electronic static dielectric constant identifies acluster involving several perovskite-type oxides and balancing with the bandgapnear the Pareto front. The method enables analyses with different viewpoints fromthose of the conventional clustering and feature importance analyses by taking therelationship between the target property and the basic features into account.RESEARCH ARTICLEwww.advintellsyst.comAdv. Intell. Syst. 2024, 2400253 2400253 (1 of 10) © 2024 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbHmailto:nobuya.sato.000@gmail.commailto:takahashi.a.bb@m.titech.ac.jphttps://doi.org/10.1002/aisy.202400253http://creativecommons.org/licenses/by/4.0/http://creativecommons.org/licenses/by/4.0/http://www.advintellsyst.comhttp://crossmark.crossref.org/dialog/?doi=10.1002%2Faisy.202400253&domain=pdf&date_stamp=2024-08-05of features in random forest or gradient boosting decision treeregression,[8–10] Shapley additive explanations,[9,11–14] variableimportance in projection scores,[15] sure independence screeningand sparsifying operator analysis regression,[16,17] and multiplelinear regression with an expectationmaximization algorithm[18,19]have been employed to extract such knowledge. Recently, naturallanguage processing has enabled the extraction ofmaterials knowl-edge even from the text corpora.[20,21] The cluster analysis is alsosuch a technique aimed at categorizing data so that similar data arein the same groups. Clustering is typically utilized for categorizingmaterials based on their basic features, such as constituentchemical elements,[17,22,23] crystal structures,[24,25] and localstructures.[26,27] The obtained groups summarize similaritiesbetween the input data points, giving us insight into their relation-ship and routes to further analysis within each group.The conventional cluster analysis is an unsupervised learningtechnique, where the target property to predict does not exist incontrast to supervised learning, such as regression. However,grading materials according to a target property as a functionof the features is often desired. For example, there is a case wherewe try to categorize semiconductors and insulators according tothe width of the bandgap and investigate the chemical and struc-tural characteristics of respective categories. This kind of cluster-ing requires taking the relationship between the basic featuresand the target property into account and, therefore, differs fromthe clustering with only the target property, which is agnosticabout the basic features. The analysis only with the target propertyis not straightforward because close target values in a cluster mayoriginate from different chemical and structural natures. Forexample, the clustering of the bandgap may gather materials intoa cluster where some bandgaps are determined primarily by theelectronegativity difference of the constituent elements, which isa good measure of ionicity, and others are determined mainly byfeatures relevant to covalency such as the average electronegativ-ity. They cannot be distinguished only by the target values, andseparate analyses with the basic chemical and structural featuresare required. In addition, the conventional clustering solely usingbasic features does not necessarily gather materials close to eachother in terms of the property of our interest.In this article, we propose a clustering method involving infor-mation about the target property as well as the basic features. Weinject information about the target property into the clusteringof the materials by the random forest (RF) regression.[28]Our method consists of three parts. First, an RF regression modelis trained for predicting a given target property. Second, the fea-ture vectors are transformed into “z-vectors” based on the paths inthe decision trees when making predictions with the trained RFmodel. Finally, the cluster analysis is performed for the z-vectors.The proposed method in this work is inspired by the past workby Breiman[29] on the analysis of the RF classification model,which divides the classes further by clustering with a data prox-imity obtained from the model. Our approach differs from such aconventional method in the following two aspects: 1) it is utilizedfor regression rather than classification problems and 2) not onlycosine similarity but also any arbitrary similarity measure can beapplied for clustering, the details of which are described later.Our method is also somewhat similar to several existing methodsto extract features from the structures of machine learning mod-els. For example, there exists the image classification method,where an ensemblemodel of classification trees is applied to trans-form an image into a vector for a subsequent classification.[30]Another example is the technique called feature learning or repre-sentation learning.[31,32] These works are similar to our method inthat data are transformed into a different form beforehand. Themain difference is that our purpose is to introduce informationabout the target variable into the clustering rather than findinga transformation specific to the subsequent task.2. Results and Discussion2.1. Comparison Between Different Target PropertiesFirst, we compare the clustering of the same dataset between twodifferent target properties, i.e., the cohesive energy per atom(Ecoh) and the bandgap (Eg) from first-principles calculations,to confirm that the results of clustering are actually differentand analyze how they differ. We applied the developed clusteringmethod to the dataset which consists of 7981 oxides collectedfrom the Materials Project database.[33,34] The distributions ofthe target properties are shown in Figure S1, SupportingInformation. Note that Eg of this dataset tends to be underesti-mated compared to experimental values; see the ComputationalSection for details. We constructed RF models using feature val-ues shown in Table S1, Supporting Information for Ecoh and Eg.The RF models are confirmed to be constructed accurately(Figure S2 and Table S2 and S3, Supporting Information).Applying the transformation into z-vectors and with the agglom-erative hierarchical clustering, we obtained dendrograms for Ecohand Eg, respectively (Figure S3 and S4, Supporting Information).In this study, we set the threshold of the distance in the z-space todivide the Eg and Ecoh datasets into 30 groups, respectively.It should be noted that appropriate thresholds of the distancein the z-space can be applied to such dendrograms to dividethe materials into any number of clusters. The higher (lower)the threshold, the finer (coarser) the classification becomes.This threshold, and therefore the number of clusters, can beadjusted to an appropriate cluster size.The obtained clusters are numbered in ascending order by themedian of the target values (Figure S3, and S4, SupportingInformation). As shown in Figure 1, the target values of eachcluster are distributed within narrower ranges than the wholedataset, which suggests that information about the target prop-erties is injected into the clustering as desired. Furthermore,there are clusters distributed in nearly the same range of the tar-get value, suggesting that the dataset is divided by not only thetarget property but also the features as desired. A comparisonbetween clusters of close target values is given later.The clusters are separated better in the target property spacethan those by the conventional clustering method simplyappending target properties to feature vectors. Figure S10,Supporting Information shows distributions of target values ofclusters obtained by the conventional clustering, i.e., the agglom-erative hierarchical clustering of the features and target variable(the number of features plus one variable), with the Euclideandistance and the Ward method. Note that the conventional clus-tering with the average method results in quite a few large clus-ters and many small clusters, which is the chaining effect, whilewww.advancedsciencenews.com www.advintellsyst.comAdv. Intell. Syst. 2024, 2400253 2400253 (2 of 10) © 2024 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbH 26404567, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aisy.202400253 by National Institute For, Wiley Online Library on [09/08/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons Licensehttp://www.advancedsciencenews.comhttp://www.advintellsyst.comour clustering does not exhibit such effects, as can be seen inFigure 1. The goodness of separation of clusters with respectto the target property is measured by the Calinski–Harabaszindex,[35] which is defined as a ratio of the intercluster dispersionto the intracluster dispersion and the larger is the better. TheCalinski–Harabasz indices of our method are 1180 for Ecohand 508 for Eg, while those of the conventional method are252 for Ecoh and 90 for Eg. The clusters of our method are sepa-rated better than those of the conventional method in terms ofthe target values because the target property is only one amonghundreds of variables in the conventional method.Although our dataset contains multication oxides, the distribu-tion of the binary oxides in respective clusters would be helpfulinformation to understand the chemical tendency. Figure 2shows which cluster each binary oxide belongs to; tabular formsummaries are given in Table S4a and S5a, SupportingInformation. As desired, the belonging of binary oxides dependson the target property. The difference is especially clear in theoxides of group 1 and 2 elements: the group 1 and 2 oxides otherthan Li2O and BeO belong to different clusters for Ecoh and thesame cluster for Eg. The difference would be related to outliers.The formation energies of group 1 oxides (excluding Li2O) areconcentrated in a low range. Specifically, they show values of2.79–3.31 eV atom�1, which are bottom 0.4% or lower in thewhole dataset (Figure 3a). The Eg value of BeO is also an outlier,which is 7.46 eV, the largest in the whole dataset (Figure 3b). Thelower Ecoh values of the group 1 oxides result partly from lowerMadelung energies due to the smaller valence of cations. In con-trast, the Eg values of the group 1 and 2 oxides are both related tooxygen p and cation s orbital characteristics in the valence andconduction bands, respectively. The distributions of their Eg val-ues are almost overlapped, which would have led to clusteringinto the same cluster.2.2. Detailed Analysis of the BandgapWe now analyze the characteristics of several clusters in theresults of the Eg clustering to see how they are related to physicaland chemical pictures. By inspecting the feature values within acluster, we can reveal the chemical and structural tendencies ofthe cluster.First, we take a closer look at cluster 3, which is the third lowestin the median of Eg among the 30 clusters. This cluster includesthe binary oxides of Cu(I) and Ag(III); namely, Cu2O and Ag2O3,and 82% of the oxides in this cluster involve at least one of Cu andAg. Although this proportion is significantly high, the proportionof such oxides of cluster 3 in the entire dataset including the otherclusters is only 34%. This suggests that cluster 3 is not classifiedsolely based on elemental information and, therefore, furtheranalysis is needed to characterize this cluster.As for characteristic structural features of cluster 3, two fea-tures denoted as Qmax4 and Qmax8 are concentrated in large values(Figure 4a,b). They are the maximums of the bond-orientationalorder parameters[36–38] among atoms defined asQmaxl ¼ maxi4π2lþ 1Xlm¼�l�����Xj∈AiΩij4πY lmðr ijÞ�����2 !1=2(1)(a) (b)Figure 1. Distributions of clusters obtained by the developed clustering method. a) and b) show the clustering for Ecoh and Eg, respectively. The clustersare numbered in ascending order by the median of the target values.www.advancedsciencenews.com www.advintellsyst.comAdv. Intell. Syst. 2024, 2400253 2400253 (3 of 10) © 2024 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbH 26404567, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aisy.202400253 by National Institute For, Wiley Online Library on [09/08/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons Licensehttp://www.advancedsciencenews.comhttp://www.advintellsyst.comwhere i is an index of an atom, Ai is a set of indices of atomsadjacent to the ith atom in terms of the Voronoi tessellation,Ωij is a solid angle from the ith atom subtended by the face com-mon to the Voronoi cells of the ith and jth atoms, Ylm(·) are thespherical harmonics, and r ij is a direction vector from the ithatom to the jth atom. Large Qmax4 and Qmax8 suggest that an oxidecontains a roughly octahedrally coordinated atom.[36] For exam-ple, ZnCu2O4, which has the largestQmax4 in this cluster, consistsof Cu atoms coordinated by six O atoms whose maximum dis-tance difference is 0.90 Å (Figure S5a, Supporting Information).Also, Li4Co3TeO8, which has the largestQmax8 in this cluster, con-sists of Li, Co, and Te atoms coordinated by six O atoms whosemaximum distance differences are 0.11, 0.11, and 0.00 Å, respec-tively (Figure S5b, Supporting Information). Note that these fea-tures are relatively less important for the whole dataset: theirpermutation importances are only 7% and 3% of the highest.The developed clustering has thus identified a series of oxidescharacterized by these locally important features in a narrowrange of Eg, unlike conventional feature importance analysisfor the whole dataset.(a)(b)Figure 2. Distributions of binary oxides in clusters obtained by the developed clustering. a) and b) show the clustering for Ecoh and Eg, respectively.Cations in binary oxides are depicted as colored annulus sectors with their oxidation numbers, where the colors correspond to the respective clusters.Cations of colorless sectors are contained only in oxides consisting of more than two cation species in the present dataset. The clusters are numbered inascending order by the median of the target values, where oxides consisting of multiple cations are also considered. The clusters with light-gray numbersdo not contain binary oxides.www.advancedsciencenews.com www.advintellsyst.comAdv. Intell. Syst. 2024, 2400253 2400253 (4 of 10) © 2024 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbH 26404567, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aisy.202400253 by National Institute For, Wiley Online Library on [09/08/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons Licensehttp://www.advancedsciencenews.comhttp://www.advintellsyst.comAs for the target property, the Eg of Cu2O, Ag2O3, andZnCu2O4 are related to the energy splitting of partially occupiedd states even in Cu2O with a formally Cu d10 configuration: boththe valence band maxima and conduction band minima aremainly characterized by the transition metal (Cu or Ag) d statesand O p states (Figure S6a–c, Supporting Information). Thevalence band maximum of Li4Co3TeO8 is also characterizedby Co d states and O p states, while the conduction band mini-mum is characterized by Te s states and O p states (Figure S6d,Supporting Information). The main peaks of the unoccupied Cod states in the density of states are a few eV higher in energy thanthe conduction band minimum.Our clustering method also helps us to analyze and under-stand oxides with similar target values originating from differentchemical and structural natures. For example, although clusters22 and 23 in the Eg case are distributed in Eg ranges close to eachother, whose medians are 3.2 and 3.3 eV, respectively, they areseparated from each other in quite early stage of the agglomer-ative hierarchical clustering and their contents show clearly dif-ferent tendencies in the chemical composition. All oxides ofcluster 23 contain N, while none of the oxides in cluster 22 con-tain N and 76% of the oxides contain at least one of Si, Ge, andAs. It would be useful to be able to classify substances with simi-lar physical and chemical properties by different origins like thiscase, which is difficult by directly applying clustering methods toonly the target properties.Looking more closely at cluster 23, it can be seen that the val-ues of Qmax9 are concentrated at large values compared to theentire dataset (Figure 4c). It suggests that cluster 23 consistsof nitrates because Qmax9 is large if the trigonal planar coordina-tion exists. Cluster 23 also involves compounds whose ratio of Natoms to O atoms is not 1:3 such as YNO4, which has an NO3local structure though its composition ratio seems not to be anitrate.Cluster 30 in the Eg case, which is the largest in the median ofEg among the 30 clusters, includes binary oxides of C(IV), P(V),and S(VI), that is, CO2, P2O5, and SO3. CO2 and SO3 are molec-ular crystals. Almost all oxides, 97%, of this cluster contain atleast one of B, P, and S. Oxides containing C are only 1% thoughits binary oxide is included in this cluster. In contrast, 34% of the(a) (b)Figure 3. Positions of group 1 and 2 binary oxides in the distributions. a) and b) show the distributions of Ecoh and Eg, respectively. Box plots indicate thedistributions of the whole dataset.(a) (b)(c) (d)Figure 4. Distributions of the maximum bond-orientational order parameters among atoms (Qmaxl ) in specific clusters for Eg. a) shows the Qmax4 distri-bution in cluster 3, theQmax8 distribution in cluster 3, theQmax9 distributions in clusters 22 (the left peak) and 23 (the right peak), and theQmax3 distributionin cluster 30. The histograms in the back and the lower box plots indicate the distribution for the whole dataset. Each histogram is normalized so that thearea is equal to one.www.advancedsciencenews.com www.advintellsyst.comAdv. Intell. Syst. 2024, 2400253 2400253 (5 of 10) © 2024 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbH 26404567, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aisy.202400253 by National Institute For, Wiley Online Library on [09/08/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons Licensehttp://www.advancedsciencenews.comhttp://www.advintellsyst.comoxides in this cluster contain B even though its binary oxide is notincluded.The inspection of the features shows that Qmax3 is large formost entries of cluster 30 (Figure 4d). In this dataset, Qmax3 is cor-related with the maximum tetrahedral order parameter,[39] whichmeasures the similarity of local structures to the tetrahedral coor-dination (with the Pearson correlation coefficient of 0.79), thoughthe maximum tetrahedral order parameter has been eliminatedfrom the regression by the feature selection. For example,KAl(SO4)2, which has the largest Qmax3 in this cluster, containsS atoms coordinated by four O atoms whose maximum distancedifference is 0.03 Å (Figure S5c, Supporting Information).The chemical formula and electronic structure of KAl(SO4)2imply its strong ionic character. The valence band ofKAl(SO4)2 is characterized predominantly by O p states, whilea minor hybridization of S s, S p, O s, and O p states can be foundin the conduction band (Figure S6e, Supporting Information).The ionic character of oxides in cluster 30 is also implied bythe feature indicating the average atomic radius, which showsa weak negative correlation with Eg (Figure 5). A small averageatomic radius is likely to correspond to short interatomic distan-ces even in oxides with a high ionicity, which causes strongCoulomb interactions and results in a large Eg value.2.3. Clustering of the Polycrystalline Average of the ElectronicStatic Dielectric TensorWe also analyze the polycrystalline average of the electronic con-tribution to the static dielectric tensor (εel) and extract features ofmaterials near the Pareto front in the Eg–εel space. The datasetconsists of 1301 oxides collected from the Materials Project data-base. Note that the εel values in this dataset tend to be overesti-mated compared to experimental values, as detailed in theComputational Section. The dataset is divided into 20 clustersby our method with the agglomerative hierarchical clustering,where the clusters are numbered in ascending order by themedian of εel (Figure S7a, Supporting Information), as in thecases of Ecoh and Eg. The number of clusters is reduced from30 to 20 because the dataset is smaller than that for Ecoh andFigure 5. Distribution of oxides with respect to the average atomic radiusand Eg. Blue diamonds indicate oxides included in cluster 30 for Eg.Figure 6. Distributions of oxides with respect to Eg and εel. Each panel shows a cluster composed of oxides indicated by colored diamonds. The clusteringis performed for εel. The clusters are numbered in ascending order by the median of εel.www.advancedsciencenews.com www.advintellsyst.comAdv. Intell. Syst. 2024, 2400253 2400253 (6 of 10) © 2024 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbH 26404567, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aisy.202400253 by National Institute For, Wiley Online Library on [09/08/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons Licensehttp://www.advancedsciencenews.comhttp://www.advintellsyst.comEg; we have confirmed that the clusters to which respective binaryoxides belong do not change by reducing the number of clusters,except for HfO2. The distributions of clusters with respect to thetarget property εel and the clusters to which binary oxides belongare shown in Figure S7b and Table S6, Supporting Information,respectively.Figure 6 shows distributions of clusters with respect to Eg andεel. We focus on a region where both Eg and εel values are rela-tively large because they are known to be in a trade-off relation-ship.[40] We find that cluster 19 is distributed near the Paretofront, which is depicted in Figure 7 in more detail. A binary oxidein cluster 19 is only that of W(VI) or WO3. All the other oxides incluster 19 also contain a transition-metal atom. For the regres-sion of εel, the features with the highest and second highest per-mutation importances are the mass density and the fraction oftransition metal atoms. The importance of mass density andits chemical interpretations have been revealed in our previousstudy.[8] Although all oxides in cluster 19 contain a transitionmetal atom, the correlation between the fraction of transitionmetal atoms and εel is not apparent, within the whole datasetor within this cluster (Figure 8a). Moreover, the mass densityis not clearly correlated within cluster 19 (Figure 8b). The otherfeatures selected for the regression are also hardly correlated withεel within this cluster: the feature measuring the average differ-ence in the number of filled valence s electrons between an atomand its neighbors gives the maximum magnitude of the Pearsoncorrelation coefficient with εel (0.44); the feature measuring theaverage difference in the number of empty valence s electronsgives the same correlation coefficient because these two featuresare identical for all oxides in the cluster 19. Although they seemto be moderately correlated with εel, this correlation might be dueto NbRhO4 which is largest in both the features (0.60) and εel(10.9). The correlation coefficient decreases to 0.37 if this oxideis omitted. The weak correlations to the features imply that thetrade-off relationship within cluster 19 is related to multiple fea-tures complexly.The feature measuring the maximum similarity of local struc-tures to the octahedral coordination[41] is concentrated in largevalues for cluster 19 (Figure 9), though the feature is eliminatedby the feature selection performed during the construction of theRF model. We have investigated all the structures in the clusterand found that they contain a metal atom octahedrallycoordinated by O atoms. Notable entries in this cluster areperovskite-type oxides, some of which are known as ferroelectriccompounds.[42] There are eight in this cluster: tetragonal SrTiO3,rhombohedral BaTiO3, tetragonal PbTiO3, tetragonal PbVO3,rhombohedral BiFeO3, orthorhombic NaTaO3, cubic KTaO3,and rhombohedral AgTaO3. The perovskite-type oxides have aFigure 7. Detail of cluster 19 obtained by the clustering for εel. The stepline shows the Pareto front. The diamonds indicate the oxides incluster 19.(a) (b)Figure 8. Distribution of oxides with respect to the εel and a) fraction of transition metal atoms (the most important feature) and b) mass density(the second most important feature). The importance of feature is ranked by random forest permutation importance. The diamonds indicate the oxidesincluded in cluster 19 for εel.Figure 9. Distribution of the maximum similarity to the octahedral coor-dination in cluster 19 for εel. The gray histogram and the lower box plotindicates the distribution for the whole dataset. Each histogram is normal-ized so that the area is equal to one.www.advancedsciencenews.com www.advintellsyst.comAdv. Intell. Syst. 2024, 2400253 2400253 (7 of 10) © 2024 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbH 26404567, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aisy.202400253 by National Institute For, Wiley Online Library on [09/08/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons Licensehttp://www.advancedsciencenews.comhttp://www.advintellsyst.comchemical composition of ABO3 and consist of corner-sharingBO6 octahedra and A atoms surrounded by eight BO6 octahedra(Figure S8a–h, Supporting Information). The ReO3-type structuretaken by WO3 is similar to the perovskite-type structure, whichconsists of corner-sharing BO6 octahedra without A atoms(Figure S8i, Supporting Information). Note that the octahedraof tetragonal PbTiO3 and tetragonal PbVO3 are significantly dis-torted and exhibit small values of the feature for the maximumsimilarity of local structures to the octahedral coordination.BaTiO3 is a transition metal oxide such that the d states of Tiare formally empty: the valence and conduction bands are domi-nated by O p states and Ti d states, respectively (Figure S9a,Supporting Information). The density of states is high and steeparound both the valence band maximum and conduction bandminimum, which complements an increase in the εel with manyelectronic transition routes from the valence to the conductionband states.[40] Consequently, both Eg and εel are relatively large.Although PbVO3 also shows large Eg and εel, it is located ratherfar from the Pareto front. As well as BaTiO3, the density of statesof PbVO3 is high and steep around the band edges (Figure S9b,Supporting Information). The difference is that the d states of Vare partially filled: the conduction band is dominated by V d states,while the valence band is characterized by O p states and a com-parable contribution of V d states. Since the d-d electronic tran-sitions are dipole-forbidden, this band structure makes εel slightlysmaller than that of BaTiO3 despite a narrower Eg.[40]3. ConclusionWe have developed a clustering method involving the informa-tion about the target property, where the information is injectedthrough a transformation of the features by the RF regressionmodel. A comparison between the clustering for Ecoh and Egdemonstrates injecting the target property information. An anal-ysis of a narrow-Eg cluster has revealed that features that are char-acteristic of each cluster are reasonable from the perspective ofconventional physical and chemical pictures, but they are notnecessarily important for the whole dataset. We have also ana-lyzed a cluster near the Pareto front in the Eg–εel space. The clus-ter consists mainly of transition metal oxides, and they show acommon structural characteristic that a metal atom is octahe-drally coordinated by O atoms. The cluster includes severalperovskite-type oxides, the electronic structures of which canexplain a balance of relatively wide Eg and large εel. Our methodenables analyses from viewpoints that are different from the con-ventional clustering and feature importance analyses by takingthe relationship between the target property and the features intoaccount. While we focus on single target property cases in thisarticle for conciseness, our method is extendable to clusteringwith respect to multiple target properties: we can constructthe RF model for each target variable, concatenate the vectorstransformed by these models, and perform the cluster analysisfor the concatenated vector.4. Computational SectionClustering Assisted by the Random Forest: Our method consists of con-structing a RF model for predicting the target property, transforming thefeatures using the model, and clustering the transformed data. A focalpoint is how the feature vectors are transformed. A schematic view isshown in Figure 10.The RF regression model consists of a set of decision trees. The deci-sion tree divides the feature space into two regions recursively, decompos-ing the feature space into an irregular grid (Figure 10a). The separation isconducted so that training data within the same region in the feature spacehave close values of the target variable. The training set of the decision treeis a bootstrap sample, i.e., a random sample from the training set of the RFmodel with replacement. Each decision tree in the RF model differentlyseparates the feature space due to randomness in the training set anddivided features. The RF regression model predicts a target value, y,for a feature vector, x, by averaging predictions by the decision treesyðxÞ ¼ 1TXTt¼1ytðxÞ (2)Figure 10. Schematic views of the clustering assisted by the RF regression. a) Region of the feature space to which a feature vector, x, belongs. For the tthdecision tree of the RF model, if the j1th component of x, xj1 , is less than or equal to θ1 and xj2 is greater than θ2, x belongs to the region labeled B, hencertðxÞ ¼ B. b) Transformation of x. For the tth decision tree, the region to which x belongs is represented by zt whose component is one if corresponding tort(x) and zero otherwise. The transformed vector, z, is a concatenation of all zt. c) Clustering in the z-space. In the feature space, x belongs to the regionslabeled r1(x), r2(x), r3(x), and so on. The more (less) regions a data point is contained in, the more similar (dissimilar) to x the data point is. For example,x 0 is more similar to x than x 00. The relationship between the number of the containing regions and the similarity (dissimilarity) measure is defined as acloseness (distance) in the z-space. A set of z is divided into clusters so that close ones are in the same clusters and distant ones are in different clusters.www.advancedsciencenews.com www.advintellsyst.comAdv. Intell. Syst. 2024, 2400253 2400253 (8 of 10) © 2024 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbH 26404567, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aisy.202400253 by National Institute For, Wiley Online Library on [09/08/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons Licensehttp://www.advancedsciencenews.comhttp://www.advintellsyst.comwhere T is the number of decision trees and yt(x) is a target value predictedby the tth decision tree. The prediction by the decision tree is constantwithin each separated region, which is an average target value of the train-ing data belonging to the regionytðxÞ ¼1jΛtrtðxÞjXi∈Λtrt ðxÞyi (3)where rt(x) is a label of the region of the tth decision tree to which xbelongs, Λtr is a bag of indices of training data used by the tth decisiontree and belonging to the region labeled r, and yi is a target value of the ithdata point in the training set of the RF model.Using the RF model, x can be nonlinearly transformed into a vector ofthe region labels to which it belongs: [r1(x), …, rT(x)]T. This categorical vec-tor can be converted into a binary vector by the one-hot encoding. Wedenote the binary vector by z(x)zðxÞ ¼ z1ðxÞ⊕ : : :⊕zT xð Þ (4)where zt(x) is a one-hot representation of rt(x), that is, the number ofdimensions is equal to the number of leaf nodes of the tth decision tree,and a component is one if corresponding to rt(x) and zero otherwise(Figure 10b). The numbers of leaf nodes of the decision trees are deter-mined by hyperparameters of the RF model, which would be tuned to giveaccurate predictions by the RF model. Note that information about thetarget variable is involved in z: the transformation from x requires rt(x),hence a regression of the target variable.The clustering is performed in the z-space. The clustering method canbe anything applicable to binary vectors. The clustering divides a set ofdata points so that similar ones are in the same clusters and dissimilarones are in different clusters (Figure 10c). The (dis)similarity measuredepends on the clustering method: the Euclidean distance is one of them.The clustering in the z-space gathers data points similar in the target vari-able as well as the features because the RF model is trained so that data inthe same region has close target values. The (dis)similarities among z canbe different even if those among x are the same. It is a consequence of thenonlinear transformation or involving the target variable.The clustering in the z-space is also encouraged because the cosinesimilarity between z is identical to the proximity measure defined inref. [29]. The proximity measure is inherent in a trained RF model, andthat between two feature vectors is defined as the proportion of decisiontrees at which the feature vectors belong to the same region:[29] it is writtenfor the feature vectors of x and x 0 asPTt¼1 I½rtðxÞ ¼ rtðx 0 Þ�=T , where I[·] isthe indicator function. Using z, the proximity can be rewritten aszðxÞTzðx0 Þ=T , which is the cosine similarity between z(x) and z(x 0). In otherwords, the clustering in the z-space is a generalization of the clusteringdescribed in ref. [29]. If the RF model is for the classification problemand the clustering in the z-space is performed with the cosine similarityand it is equivalent to the method of ref. [29].In our demonstrations for different target properties, the RF regressionmodel is constructed for each target property by the scikit-learn library.[43]Hyperparameters are tuned by the fivefold cross-validation. The number offeatures is reduced by the recursive feature elimination:[44] the number isdetermined by the fivefold cross-validation using the tuned hyperpara-meters. The features are ranked by the permutation feature importance,[28]and those in the lowest 20% are removed recursively for each fold. Themodel used for the clustering is trained using the whole dataset withthe tuned hyperparameters and the selected number features.Reducing the number of features facilitates the investigation into clustersconcerning the features. The εel values are transformed by the commonlogarithm beforehand because large numerical errors are expected forlarge εel values.The clustering in the z-space is performed by the agglomerative hierar-chical clustering implemented in the SciPy library.[45] The agglomerativehierarchical clustering recursively merges a closest pair of clusters intoa new cluster, where a data point is treated as a cluster consisting onlyof the data point. The closest pair is determined by the distances betweendata points. The way of determining the closest pair is called linkage cri-terion. In short, the method of the agglomerative hierarchical clustering isspecified by the distance metric for data points and the linkage criterion.We perform the clustering with four methods: combinations of two dis-tance metrics and two linkage criteria. We use the cosine distance andJaccard distance as the distance metric, and the average method (theunweighted pair group method with arithmetic mean), and the Wardmethod as the linkage criteria. We show only results with the combinationof the cosine distance and the averagemethod in the main text because wefind that the four combinations end in similar groupings of binary oxides.Results with the other combinations are presented in the SupportingInformation.Datasets: We prepare two datasets for different target properties: onefor Ecoh and Eg, and the other for εel. Both datasets are collected from theMaterials Project database,[33,34] a collection of properties computedbased on density functional theory with the Perdew–Burke–Ernzerhofparametrization of the generalized gradient approximation[46] andHubbard U corrections.[47] The dataset for Ecoh and Eg consists of 7981compounds satisfying the following conditions: 1) O atoms are contained,2) H and noble gas atoms are not contained, 3) anions are only O2�, 4) thetotal energy is the lowest among polymorphs, 5) the formation energy isless than 0.1 eV atom�1 against that of a mixture of competing phases,6) Eg is determined, and 7) Eg is larger than or equal to 0.2 eV.Oxidation states of atoms (ions) are determined by the pymatgen library[48]based on given chemical compositions. The dataset for εel consists of 1301compounds satisfying the following conditions in addition to above (1–7):(8) the static dielectric tensor is computed, and (9) εel is less than 50. Notethat the collected Eg and εel values tend to be underestimated[49,50] andoverestimated[49] compared to experimental values, respectively, owingto the generalized gradient approximation, especially for compoundstreated without the Hubbard U corrections. The distributions of the targetproperties are shown in Figure S1, Supporting Information.We generate features of oxides from crystal structures by the matminerlibrary[38] (see Table S1, Supporting Information for the generated fea-tures). After omitting nonnumerical features, constant features, and fea-tures not available for all oxides in the datasets, we obtain 696 features forEcoh and Eg, and 662 features for εel.Supporting InformationSupporting Information is available from the Wiley Online Library or fromthe author.AcknowledgementsThis work was supported by JST CREST grant no. JPMJCR17J2, JSPSKAKENHI grant no. JP20H00302, and 21K14401, MEXT Data Creationand Utilization Type Material Research and Development Projectgrant no. JPMXP1122683430, MEXT Design and Engineering byJoint Inverse Innovation for Materials Architecture Project, andKISTEC Project. The computing resource of the Academic Center forComputing and Media Studies at Kyoto University was used for part ofthis work.Conflict of InterestThe authors declare no conflict of interest.Data Availability StatementThe data that support the findings of this study are openly available inMaterials project database at https://legacy.materialsproject.org, refer-ence number [31]. The underlying code for this work is available athttps://github.com/nbsato/forestcluster.www.advancedsciencenews.com www.advintellsyst.comAdv. Intell. Syst. 2024, 2400253 2400253 (9 of 10) © 2024 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbH 26404567, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aisy.202400253 by National Institute For, Wiley Online Library on [09/08/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons Licensehttps://github.com/nbsato/forestclusterhttp://www.advancedsciencenews.comhttp://www.advintellsyst.comKeywordsclustering, inorganic compounds, interpretable artificial intelligence,random forestReceived: March 28, 2024Revised: July 1, 2024Published online:[1] K. Rajan, Mater. Today 2005, 8, 38.[2] K. T. Butler, D. W. Davies, H. Cartwright, O. Isayev, A. Walsh, Nature2018, 559, 547.[3] W. Sha, Y. Guo, Q. Yuan, S. Tang, X. Zhang, S. Lu, X. Guo, Y.-C. Cao,S. Cheng, Adv. Intell. Syst. 2020, 2, 1900143.[4] C. Yan, G. Li, Adv. Intell. Syst. 2023, 5, 2200243.[5] F. Oviedo, J. L. Ferres, T. Buonassisi, K. T. Butler, Acc. Mater. Res.2022, 3, 597.[6] G. Pilania, Comput. Mater. Sci. 2021, 193, 110360.[7] X. Zhong, B. Gallagher, S. Liu, B. Kailkhura, A. Hiszpanski,T. Y.-J. Han, NPJ Comput. Mater. 2022, 8, 204.[8] A. Takahashi, Y. Kumagai, J. Miyamoto, Y. Mochizuki, F. Oba, Phys.Rev. Mater. 2020, 4, 103801.[9] N. T. P. Hartono, J. Thapa, A. Tiihonen, F. Oviedo, C. Batali, J. J. Yoo,Z. Liu, R. Li, D. F. Marrón, M. G. Bawendi, T. Buonassisi, S. Sun,Nat.Commun. 2020, 11, 4172.[10] K. Choudhary, K. F. Garrity, V. Sharma, A. J. Biacchi, A. R. HightWalker, F. Tavazza, NPJ Comput. Mater. 2020, 6, 64.[11] S. M. Lundberg, S.-I. Lee, in Advances in Neural Information ProcessingSystems, Curran Associates, Inc., Glasgow, Scotland 2017.[12] K. Morita, D. W. Davies, K. T. Butler, A. Walsh, J. Chem. Phys. 2020,153, 024503.[13] S. Fujii, Y. Shimizu, J. Hyodo, A. Kuwabara, Y. Yamazaki, Adv. EnergyMater. 2023, 13, 2301892.[14] T. A. R. Purcell, M. Scheffler, L. M. Ghiringhelli, C. Carbogno, NPJComput. Mater. 2023, 9, 112.[15] Y. Noda, M. Otake, M. Nakayama, Sci. Technol. Adv. Mater. 2020,21, 92.[16] R. Ouyang, S. Curtarolo, E. Ahmetcik, M. Scheffler, L. M. Ghiringhelli,Phys. Rev. Mater. 2018, 2, 083802.[17] L. Sbailò, Á. Fekete, L. M. Ghiringhelli, M. Scheffler, NPJ Comput.Mater. 2022, 8, 250.[18] F. R. Burden, D. A. Winkler, QSAR Comb. Sci. 2009, 28, 645.[19] P. Mikulskis, M. R. Alexander, D. A. Winkler, Adv. Intell. Syst. 2019, 1,1900045.[20] V. Tshitoyan, J. Dagdelen, L. Weston, A. Dunn, Z. Rong, O. Kononova,K. A. Persson, G. Ceder, A. Jain, Nature 2019, 571, 95.[21] T. Gupta, M. Zaki, N. M. A. Krishnan, Mausam, NPJ Comput. Mater.2022, 8, 102.[22] I. E. Castelli, K. W. Jacobsen, Model. Simul. Mater. Sci. Eng. 2014, 22,055007.[23] W. Sun, C. J. Bartel, E. Arca, S. R. Bauers, B. Matthews, B. Orvañanos,B.-R. Chen, M. F. Toney, L. T. Schelhas, W. Tumas, J. Tate,A. Zakutayev, S. Lany, A. M. Holder, G. Ceder, Nat. Mater. 2019,18, 732.[24] E. L. Willighagen, R. Wehrens, P. Verwer, R. de Gelder,L. M. C. Buydens, Acta Crystallogr. B 2005, 61, 29.[25] B. Meredig, C. Wolverton, Chem. Mater. 2014, 26, 1985.[26] T. L. Pham, H. Kino, K. Terakura, T. Miyake, H. C. Dam, J. Chem. Phys.2016, 145, 154103.[27] R. Tamura, M. Matsuda, J. Lin, Y. Futamura, T. Sakurai, T. Miyazaki,Phys. Rev. B 2022, 105, 075107.[28] L. Breiman, Mach. Learn. 2001, 45, 5.[29] L. Breiman, https://www.stat.berkeley.edu/users/breiman/wald2002-2.pdf (accessed: January 2023).[30] F. Moosmann, B. Triggs, F. Jurie, in Advances in Neural InformationProcessing Systems, MIT Press, Cambridge, MA 2007.[31] I. Goodfellow, Y. Bengio, A. Courville, Deep Learning, MIT Press,Cambridge, MA 2016.[32] G. Zhong, L.-N. Wang, X. Ling, J. Dong, J. Finance Data Sci. 2016, 2,265.[33] A. Jain, S. P. Ong, G. Hautier, W. Chen, W. D. Richards, S. Dacek,S. Cholia, D. Gunter, D. Skinner, G. Ceder, K. A. Persson, APLMater. 2013, 1, 011002.[34] S. P. Ong, S. Cholia, A. Jain, M. Brafman, D. Gunter, G. Ceder,K. A. Persson, Comput. Mater. Sci. 2015, 97, 209.[35] T. Calinski, J. Harabasz, Commun. Stat. Theory Methods 1974, 3, 1.[36] P. J. Steinhardt, D. R. Nelson, M. Ronchetti, Phys. Rev. B 1983, 28,784.[37] A. Seko, H. Hayashi, K. Nakayama, A. Takahashi, I. Tanaka, Phys. Rev.B 2017, 95, 144110.[38] L. Ward, A. Dunn, A. Faghaninia, N. E. R. Zimmermann, S. Bajaj,Q. Wang, J. Montoya, J. Chen, K. Bystrom, M. Dylla, K. Chard,M. Asta, K. A. Persson, G. J. Snyder, I. Foster, A. Jain, Comput.Mater. Sci. 2018, 152, 60.[39] N. E. R. Zimmermann, A. Jain 2017, in progress.[40] F. Naccarato, F. Ricci, J. Suntivich, G. Hautier, L. Wirtz,G.-M. Rignanese, Phys. Rev. Mater. 2019, 3, 044602.[41] N. E. R. Zimmermann, M. K. Horton, A. Jain, M. Haranczyk, Front.Mater. 2017, 4, 34.[42] R. E. Cohen, Nature 1992, 358, 136.[43] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion,O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg,J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot,É. Duchesnay, J. Mach. Learn. Res. 2011, 12, 2825.[44] I. Guyon, J. Weston, S. Barnhill, V. Vapnik, Mach. Learn. 2002,46, 389.[45] P. Virtanen, R. Gommers, T. E. Oliphant, M. Haberland, T. Reddy,D. Cournapeau, E. Burovski, P. Peterson, W. Weckesser, J. Bright,S. J. van der Walt, M. Brett, J. Wilson, K. J. Millman, N. Mayorov,A. R. J. Nelson, E. Jones, R. Kern, E. Larson, C. J. Carey, İ. Polat,Y. Feng, E. W. Moore, J. VanderPlas, D. Laxalde, J. Perktold,R. Cimrman, I. Henriksen, E. A. Quintero, C. R. Harris, et al.,Nat. Methods 2020, 17, 261.[46] J. P. Perdew, K. Burke, M. Ernzerhof, Phys. Rev. Lett. 1996, 77,3865.[47] V. I. Anisimov, J. Zaanen, O. K. Andersen, Phys. Rev. B 1991, 44, 943.[48] S. P. Ong, W. D. Richards, A. Jain, G. Hautier, M. Kocher, S. Cholia,D. Gunter, V. L. Chevrier, K. A. Persson, G. Ceder, Comput. Mater. Sci.2013, 68, 314.[49] F. Oba, Y. Kumagai, Appl. Phys. Express 2018, 11, 060101.[50] Y. Hinuma, A. Grüneis, G. Kresse, F. Oba, Phys. Rev. B 2014, 90,155405.www.advancedsciencenews.com www.advintellsyst.comAdv. Intell. Syst. 2024, 2400253 2400253 (10 of 10) © 2024 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbH 26404567, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1002/aisy.202400253 by National Institute For, Wiley Online Library on [09/08/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons Licensehttps://www.stat.berkeley.edu/users/breiman/wald2002-2.pdfhttps://www.stat.berkeley.edu/users/breiman/wald2002-2.pdfhttp://www.advancedsciencenews.comhttp://www.advintellsyst.com Target Material Property-Dependent Cluster Analysis of Inorganic Compounds 1. Introduction 2. Results and Discussion 2.1. Comparison Between Different Target Properties 2.2. Detailed Analysis of the Bandgap 2.3. Clustering of the Polycrystalline Average of the Electronic Static Dielectric Tensor 3. Conclusion 4. Computational Section