# Fileset

[s44172-024-00227-y.pdf](https://mdr.nims.go.jp/filesets/0fabc8ab-363b-4a9f-a4ee-137c9ca3a74b/download)

## Creator

[Daiki Nishioka](https://orcid.org/0000-0002-3369-7700), [Takashi Tsuchiya](https://orcid.org/0000-0002-6950-6160), [Masataka Imura](https://orcid.org/0000-0002-4236-9549), [Yasuo Koide](https://orcid.org/0000-0001-8321-9822), Tohru Higuchi, [Kazuya Terabe](https://orcid.org/0000-0003-3988-3456)

## Rights

[Creative Commons BY Attribution 4.0 International](https://creativecommons.org/licenses/by/4.0/)

## Other metadata

[A high-performance deep reservoir computer experimentally demonstrated with ion-gating reservoirs](https://mdr.nims.go.jp/datasets/1ad07945-da9e-4e80-bc38-83ff951600c0)

## Fulltext

A high-performance deep reservoir computer experimentally demonstrated with ion-gating reservoirscommunications engineering Articlehttps://doi.org/10.1038/s44172-024-00227-yA high-performance deep reservoircomputer experimentally demonstratedwith ion-gating reservoirsCheck for updatesDaiki Nishioka 1,2, Takashi Tsuchiya 1 , Masataka Imura 3, Yasuo Koide 4, Tohru Higuchi 2 &Kazuya Terabe1While physical reservoir computing is a promising way to achieve low power consumptionneuromorphic computing, its computational performance is still insufficient at a practical level. Onepromising approach to improving its performance is deep reservoir computing, in which thecomponent reservoirs are multi-layered. However, all of the deep-reservoir schemes reported so farhave been effective only for simulation reservoirs and limited physical reservoirs, and there have beenno reports of nanodevice implementations. Here, as an ionics-based neuromorphic nanodeviceimplementation of deep-reservoir computing, we report a demonstration of deep physical reservoircomputing with maximum of four layers using an ion gating reservoir, which is a small and high-performance physical reservoir.While the previously reported deep-reservoir scheme did not improvethe performance of the ion gating reservoir, our deep-ion gating reservoir achieved a normalizedmeansquarederror of 9.08×10−3 on a second-order nonlinear autoregressivemoving average task,which isthe best performance of any physical reservoir so far reported in this task.More importantly, the deviceoutperformed full simulation reservoir computing. The dramatic performance improvement of the iongating reservoir with our deep-reservoir computing architecture paves the way for high-performance,large-scale, physical neural network devices.Physical reservoir computing (PRC), which directly utilizes the nonlineardynamics inherent in physical systems for information processing, hasattracted attention in recent years because it can drastically reduce thecomputational resources required for information processing1,2. Non-linearity, short-term memory, and high dimensionality are required forreservoirs that map input data nonlinearly into a high-dimensional fea-ture space3,4, and physical devices with these features are promising forPRC1. Many types of physical reservoirs have been reported, includingmemristors, optical devices, spintronics devices, soft bodies, nanowirenetworks, and ion-gating reservoirs. Further, information processing,including image recognition, spoken digit recognition, and time seriesprediction tasks, have been demonstrated using such physical reservoirdevices1,2,5–27. The efficiency of PRCbased onmaterial-based computationmakes it particularly promising for application to resource-limited edgeAI devices. However, the computational performance of PRC is stillinsufficient for practical information processing tasks performed by suchmaterial-based efficient edge AI devices (e.g., dynamic image recognitionand time series analysis tasks for in-situ processing of time series outputfrom sensors, such as blood glucose level prediction and speech sentencerecognition), an approach is required in order to substantially improveperformance. One way to achieve this is deep reservoir computing (deep-RC), in which reservoirs aremultilayered. Thismethod is expected to be apromising approach, just as neural networks (NN) have been shown toachieve high expressive power and performance by utilizing deep layer-ing. In full-simulation reservoir computing (RC), multilayering of thereservoir parts has been considered, and it has been reported that layeringsmall reservoirs can improve performance in comparison to single-layered reservoirs with the same total reservoir size28–33. Furthermore, ithas been reported that not only deepening the reservoir layer but alsotraining the connection weights between layers according to the taskimproves the expressiveness of the network and the flexibility of themodel, as well as the computational performance and efficiency29,30.1Research Center for Materials Nanoarchitectonics (MANA), National Institute for Materials Science (NIMS), 1-1 Namiki, Tsukuba, Ibaraki 305-0044, Japan.2Department of Applied Physics, Faculty of Science, Tokyo University of Science, Katsushika, Tokyo 125-8585, Japan. 3Research Center for Functional Materials,NIMS, 1-1 Namiki, Tsukuba, Ibaraki 305-0044, Japan. 4Research Network and Facility Services Division, NIMS, 1-2-1 Sengen, Tsukuba, Ibaraki 305-0047, Japan.e-mail: TSUCHIYA.Takashi@nims.go.jpCommunications Engineering |            (2024) 3:81 11234567890():,;1234567890():,;http://crossmark.crossref.org/dialog/?doi=10.1038/s44172-024-00227-y&domain=pdfhttp://crossmark.crossref.org/dialog/?doi=10.1038/s44172-024-00227-y&domain=pdfhttp://crossmark.crossref.org/dialog/?doi=10.1038/s44172-024-00227-y&domain=pdfhttp://orcid.org/0000-0002-3369-7700http://orcid.org/0000-0002-3369-7700http://orcid.org/0000-0002-3369-7700http://orcid.org/0000-0002-3369-7700http://orcid.org/0000-0002-3369-7700http://orcid.org/0000-0002-6950-6160http://orcid.org/0000-0002-6950-6160http://orcid.org/0000-0002-6950-6160http://orcid.org/0000-0002-6950-6160http://orcid.org/0000-0002-6950-6160http://orcid.org/0000-0002-4236-9549http://orcid.org/0000-0002-4236-9549http://orcid.org/0000-0002-4236-9549http://orcid.org/0000-0002-4236-9549http://orcid.org/0000-0002-4236-9549http://orcid.org/0000-0001-8321-9822http://orcid.org/0000-0001-8321-9822http://orcid.org/0000-0001-8321-9822http://orcid.org/0000-0001-8321-9822http://orcid.org/0000-0001-8321-9822http://orcid.org/0000-0002-0491-5545http://orcid.org/0000-0002-0491-5545http://orcid.org/0000-0002-0491-5545http://orcid.org/0000-0002-0491-5545http://orcid.org/0000-0002-0491-5545mailto:TSUCHIYA.Takashi@nims.go.jpIn particular, Deep-Echo State Network, with hyperbolic tangents as anonlinear function, has shown remarkable performance improvementsin the prediction tasks of nonlinear autoregressive moving averagemodels and chaotic dynamical systems when the connection weightsbetween the layers are trained by linear regressionwith the targets29,30. Onthe other hand, few attempts at multilayering in physical reservoirs orphysical NN have been reported, and those that have been reported arelimited to methods that either do not train the connection weightsbetween reservoirs (i.e., the network is not highly flexible)34,35 or train theconnection weights between reservoirs (or layers of physical NN) usingbackpropagation algorithms that require complex calculations that relyon external circuits and have large computational costs33,36. It is parti-cularly notable that there are no reports of deep-RCwith nanodevices thatare advantageous for integration to realize practical AI devices, and thus,it is not clear that multilayering is effective in improving the performanceof physical reservoirs.Here, as the first implementation of nanodevice-based Deep-RC, wedescribe a demonstrationof deepphysical reservoir computingusing an ion-gating reservoir (IGR), which is a compact and high-performance physicalreservoir. The IGR is a nanodevice with a transistor structure consisting of ahydrogen-terminated diamond channel and a Li+ electrolyte (Li-Si-Zr-O)23,37, and the ion-electron coupled dynamics based on the electric double-layer (EDL) effect in the nanoregion near the Li+ electrolyte/diamondinterface is used as nonlinear dynamics for the reservoir computing.There are two approaches to achieving high-performance PRC: one isto improve the performance by modifying the physical reservoir itself, andthe other is to maximize the information processing capability of the phy-sical systemthrough systemapproaches suchas optimizationof thenetworkstructure and external data manipulation (e.g., pre-processing such asmasking). In this study, we employed the latter approach, which improvesPRC performance through deep network architecture, and used the IGR,which has achieved high computational performance in time series dataanalysis and image recognition tasks23,26, as amodel device for this approach.First, we verified the computational performance of the deep-RCscheme reported for simulation-RC using IGR28–30. In this network, theconnection weights between reservoir layers are trained based on a simplelinear regression algorithm, which provides a higher network flexibilitycompared to the scheme in which the connection weights between reser-voirs are not trained34,35, anddoesnot require a back-propagation algorithm.The backpropagation algorithm is an effectivemethod that greatly improvesthe expressive power of the network, but it is difficult to apply to PRCs basedon complex and dynamic nonlinearities (black box functions) originatingfrom physical systems because it requires detailed information on thenonlinearities in the reservoir layer and their derivatives33,36. In this respect,the method of learning weights between layers by linear regression usingtargets is well suited for implementation in physical systems because it doesnot require detailed information on nonlinearities in physical systems (andtheir derivatives, etc.) anddoesnot require inverse input of errors to physicalsystems inorder tobackpropagate the errors28–30.However, the conventionalscheme with a simple series structure network reported for the simulation-RC does not improve the computational performance of the IGR. This wasfound to be due to the inherent characteristic of general physical reservoirs,which are sensitive to input conditions. In addition, this scheme does notprovide any improvement in network size limitation, which is one of themain reasons that severely inhibited the performanceof conventional PRCs.Whereas in simulation-RC it is easy to increase the reservoir size to thedesired performance in exchange for computational cost, in PRCs thenumber of reservoir states (current and voltage response, mechanicalvibration, optical response, etc.) obtained from the physical device is limitedby themeans of access to the physical system, such asmeasurement probes.Thus, it is difficult to increase the network size of such PRCs.On theother hand, a deep-RCschemewithparallel structure overcomesuch limitations of physical reservoirs in principle, and succeeded in greatlyincreasing the numberof reservoir states by utilizing the reservoir outputs ofthe previous layer as well as the final outputs among the layered deviceoutputs. As a result, the performance of the IGR applied with the subjectdeep-RC scheme (deep-IGR) was drastically improved, with a 41% reduc-tion in error compared to single-layer IGR in a second-order nonlinearautoregressive moving average (NARMA2) task. However, the model alsoshowed an increase of unnecessary reservoir states leading to overlearning.Therefore, a modified model was developed to improve this model byevaluating the high dimensionality of the reservoir state in terms of thecorrelation coefficient between reservoir states, the number of layers wasincreased by excluding featureless reservoir states (which do not contributeto high dimensionality), which resulted in a 53% reduction in error for theNARMA2 task compared to single-layer IGR, with a normalized meansquared error (NMSE) of 9.08 × 10−3. This is the best performance of anyphysical reservoir reported to date22–25,38–40, outperforming a full-simulationRC38 for the NARMA2 task.ResultsIon-gating reservoirs using electric double-layer transistorsTo demonstrate the applicability of the deep-RC architecture, we employedan IGR as the physical reservoir part, implemented with an EDL transistorcomposed of a Li+ electrolyte (Li-Si-Zr-O) and hydrogen-terminated dia-mond, as shown in Fig. 1a. The IGR transistor has 9 channels of differentlengthsLch (=5, 10, 25, 35, 50, 100, 250, 350, 500 µm), and the 9drain currentresponses can be nonlinearly transformed by the EDL mechanism to theinput gate voltage signal. Figure 1b shows the drain current (ID)-gate voltage(VG) curves obtained from channels with lengths of Lch = 50 µm andLch = 500 µm(upper panel) and agate current (IG)–VG curve (lower panel).With a positive gate voltage applied, Li+ in the electrolyte accumulates at theelectrolyte/diamond interface, and electrons are injected into the diamond,which is ahole conductor, resulting in ahigh resistance stateof thediamond.On the other hand, when a negative gate voltage is applied, negativelycharged Li vacancies in the electrolyte accumulate on the diamond surface,forming an EDL at the electrolyte/diamond interface as shown in Fig. 1a.Then, holes are doped into the hydrogen-terminated diamond channel,resulting in a low-resistivity state and the drain current is nonlinearlymodulated37. As different channel resistances exhibit different relaxationtimes for different channel lengths, if the gate input is a pulse signal, it willexhibit different drain current responses depending on the channel length,as shown in Fig. 1c, providing a higher dimension as a physical node23. Thestructure of this device, where different channel lengths coexist, providescurrent responses with different relaxation times. The coexistence in onesystem of a current response with a short relaxation time, which reflects ashort-term experience and has strong nonlinearity, and a current responsewith a long relaxation time, which reflects a long-term experience andhas relatively weak nonlinearity, can be expected to provide highperformance1,23,41. In addition to these drain currents, the gate currents,obtained from the input gate terminals, show a spiked response, as shown inthe bottom panel of Fig. 1c, which provides additional diversity to theIGR24,26. In redox-based IGR, the use of gate currents as reservoir states hasbeen reported to improve computational performance24. In addition to the10 physical nodes obtained by adding the gate current response to the 9drain current responses, 10 nodes per pulse response were obtained asvirtual nodes, as shown in the right-hand panel of Fig. 1c. Thus, the numberof reservoir statesXiobtained from the IGR (i.e., nodes) is 100. Said reservoirstatesXiwerenormalized from0 to1and subsequentlyused for the reservoirpart in the schematic of reservoir computing shown in Fig. 1d. The reservoiroutput y(k) at discrete time step k is obtained by the linear combination ofreservoir state Xi(k) and readout weightWi as follows;y kð Þ ¼ PNi¼1WiXi kð Þ þ b¼ WX kð Þ þ bð1Þwhere b is the bias; N is the reservoir size; W = (W1, W2, …, WN) is thereadout weight vector; X(k)=[X1(k), X2(k),…, XN(k)]T is the reservoir statevector. Figure 1e is a schematic of the deep-IGR which is a physicallyhttps://doi.org/10.1038/s44172-024-00227-y ArticleCommunications Engineering |            (2024) 3:81 2implemented deep reservoir computing with IGRs. In the first layer, as inconventional IGRs23, the voltage-transformed input signal was input to thegate of the IGR, and the reservoir output was calculated by a linear sum ofthe weights and reservoir states (Eq. 1) obtained by acquiring virtual nodesfrom the obtained current responses, as shown in Fig. 1c. The weights weretrained by linear regression so that the target and reservoir outputsmatched.In deep-IGR, the reservoir output of the first layer is converted to a voltagepulse and input to the gate of the second layer IGR, and the reservoir outputis obtained via weights using the obtained current response as in the firstlayer. The input that reproduced the target waveform to some degree in thefirst layer is againnonlinearly transformedby IGR into ahigherdimensionalfeature space for learning, which allows the representation of target featuresthat could not be represented in thefirst layer. Thenumber of said layers canbe increased by inputting the reservoir outputs of the previous layer, in thesame way as in the second layer, for the third and subsequent layers. Theperformance of deep-IGRwas evaluated by the error between the target andthe reservoir output obtained by the input and forward propagation of adataset different fromthe trainingdataset,with theweights of all layersfixed.The details of deep-IGR are discussed later in this document.Before evaluating the performance of deep-IGR, we performed aNARMA2 task that predicts the NARMA2model shown in Eq. 2 followinga prior study42, in order to evaluate the computational performance of thesingle-layer IGR.ytðkþ 1Þ ¼ 0:4ytðkÞ þ 0:4ytðkÞytðk� 1Þ þ 0:6u3ðkÞ þ 0:1 ð2Þwhere yt(k) are themodel outputs;u(k) are the random inputs, ranging from0 to 0.5. The NARMA2 task, which is widely used as a benchmark task forphysical reservoirs, requires reservoirs to exhibit second-order nonlinea-rities and short-term memory22–25,38–40. Figure 2a shows a schematic of theNARMA2 task performed by IGR. The random input signal u(k) wasconverted into voltagepulse streamswith an intensity of 0 V to0.5 V, a pulseperiod ofT, and a duty cycle ofD, then input to the gate terminal of the IGRtransistor. The responses of the drain current to the gate voltage pulsestreams were measured at a constant drain voltage of−0.5 V. As shown inFig. 1c, 100 reservoir statesXiwere obtained from 10 current responses andvirtual nodes (i = 1, 2, …, 100). Furthermore, an additional 100 reservoirstates Xi (i = 101, 102, …, 200) were obtained by applying an intensity-reversed input uinv(k) [= 0.5 - u(k)] to the IGR (inversion pulse method),resulting in a total of 200 reservoir states, the combination of which wasutilized to obtain the reservoir output as shown in Eq. 1. Details on theinversion pulse method and reservoir size of the single IGR are givenSupplementaryNote 1 and Supplementary Figs. S1-3. In the training phase,the readout weights were trained by linear regression, in order to match theFig. 1 | Physical reservoir computing using an iongating reservoir (IGR). a Schematic of the IGRtransistor. The inset is an optical microscope imageof a diamond with source and drain electrodes. b ID-VG curves (upper panel) for channels with lengths of50 µm and 500 µm and IG-VG curve (lower panel).c ID responses (Lch = 50 µm, 500 µm) and IGresponse to pulsed VG input. The right-hand panelshows how to obtain virtual nodes from such currentresponses. d Schematic of reservoir computing.e Schematic of the deep reservoir computingimplemented by IGRs.https://doi.org/10.1038/s44172-024-00227-y ArticleCommunications Engineering |            (2024) 3:81 3reservoir output y(k) and the model output yt(k). Details of the trainingalgorithm are given in the Method section herein. In the test phase,performance was evaluated by NMSE (Eq. 3) of the reservoir output(prediction) y(k) to the model output yt(k) for a different input u(k) than inthe training phase.NMSE ¼ 1MPMk¼1 yt kð Þ � y kð Þ� �2σ2 yt kð Þ� � ð3Þwhere M is the data length (M = 1600 for the training phase andM = 700 for the test phase); σ2(·) is the variance. Figure 2b and c showthe pulse period T and duty cycle D dependence of NMSEs in thetraining and test phases, respectively. The best results, for both trainingand test phases, were obtained at T = 70 ms and D = 70%, where theNMSE was 0.0157 in the training phase and 0.0194 in the test phase.Supplementary Fig. S4a shows the D dependence of NMSE atT = 70 ms, and Supplementary Fig. S4b shows the T dependence ofNMSE at D = 70%. Both of these are minima at the optimal conditions(indicated by *), indicating that the search for optimal conditions insingle-layer IGR was successfully performed. In other words, this is thelimit of the computational performance that can be achieved byadjusting the pulse period and duty cycle. Further performanceimprovements require consideration of voltage conditions (VG andVD) and preprocessing of the input signal (feature extraction, masking,etc.). Tuning these so-called ‘hyperparameters’ is difficult for physicalreservoirs that require actual measurements, and the huge variety ofparameters (voltage, time, temperature, number of masks, etc.) andtheir combinations make it extremely difficult to maximize thepotential computational performance of the physical system. In all ofthe deep-IGR experiments discussed below, the voltage pulseconditions were fixed at T = 70 ms and D = 70%. In this case, theoptimal conditions match between training and testing, but if such isnot the case, the optimal conditions in the training data should beadopted as the input conditions for the second and subsequent layers inorder to avoid incorrect optimization by the testing data.Performance evaluation of deep ion-gating reservoir withNARMA2 taskWe experimentally demonstrated the deep-IGR shown in Fig. 3a,b as aframework that overcomes the limitations discussed above, and easilymaximizes the computational performance of the physical system.Although the deep layered network structure shown in Fig. 3a (Network 1)has been reported to improve performance in a full simulation reservoir28–30,there are no reports of its application to a physical reservoir. In the first layerof Network 1, the readout weightW(1) is trained with the same procedure aswith the single-layer IGR shown in Fig. 2a, so as to obtain the first-layerreservoir output y(1)(k). In the L-th layerðL^2Þ, the voltage-transformedreservoir output y(L−1)(k) from the previous layer is input to the IGR insteadof random input u(k). The reservoir output of the L-th layer is then calcu-lated by the linear combination of the reservoir state X(L)(k) obtained fromthe IGR and the readout weightW(L) trained by linear regression. In the testphase, the signal was propagated forwardwith theweights of all layers fixed,and the computational performance was evaluated by the error between thereservoir output y(L)(k) and themodel output yt(k) (Eq. 2) at thefinal layerL.For details on the operating time in our Deep-RC scheme, please refer toSupplementary Note 3 and Supplementary Figs. S5 and S6. The black andgray plots in Fig. 3c show the dependence of the NMSE on the number oflayers in the test phase and the training phase, respectively, for theNARMA2 taskperformedonNetwork1 to the trainingdata. In this task, thereservoir is used to predict the NARMA2 model output yt(k) from inputu(k). However, in the structure of Network 1 (deep layered), the nature ofthe problem changes after the second layer, and the task switches to pre-dicting the NARMA2 model output using the reservoir prediction of theprevious layer as input (task switching). In a physical reservoir that usesthe transient response of a physical system in real time, the properties of thereservoir state (nonlinearity, memory capacity, and diversity) changestrongly depending on the operating conditions (in this case, the input VGpulse stream conditions), so that the accuracy in a given task changes due tosensitivity to the device operating conditions (i.e., as discussed in Fig. 2, theperformance varies greatly with the operating conditions of the IGR.).Therefore, in this case, where the nature of the task has changed, differentFig. 2 | Performance evaluation of a single iongating reservoir (IGR) by the second-order non-linear autoregressive moving average(NARMA2) task. a Schematic of the NARMA2 taskperformed by IGR, showing the pulse period T andduty cycle D dependence of normalized meansquared errors (NMSEs) for training (b) and test (c)phases. Optimal conditions are indicated by *.https://doi.org/10.1038/s44172-024-00227-y ArticleCommunications Engineering |            (2024) 3:81 4operating conditions for the IGR need to be explored because differentcharacteristics are required for the reservoir. However, the method ofsearching for optimal operating conditions for each additional layer is notrealistically achievable. In addition, although the task does not changeconsiderably after the second layer, the input to the reservoir (i.e., the outputof the previous layer) changes according to the learning status of the net-work, so optimization of the operating conditions is still considerednecessary.Therefore, in order to overcome such drawback and improve perfor-mance, without the need to search for optimal operating conditions, weconsidered Network 2 (Fig. 3b). This network is a modified model of Net-work1 (deep layered) that utilizes all the reservoir states obtained from thereservoirs of theprevious layers [X(1)(k),X(2)(k),…,X(L)(k)] so as to obtain thereservoir output of a given L-th layer (deep and parallel structure). In thisscheme, higher dimensionality (the number of reservoir states used toobtain output) increases with the number of layers, so that the expressive-ness of the reservoir also increases. Anothermajor difference fromNetwork1 (deep layered) is that task switching does not occur as discussed in Net-work 1. Since the reservoir stateX(1)(k) in the first layer is always used in thecomputation, regardless of the number of layers, the reservoir states in thesecond and subsequent layers can be interpreted as additional features thatimprove the accuracy of predicting the NARMA2model from the reservoirstates in the first layer. The blue and light blue plots show the dependence ofthe NMSE on the number of layers in the test phase and the training phase,respectively, for the NARMA2 task performed on Network 2 (deep andparallel layered)28. The NMSE decreases dramatically as the number oflayers increases, and the error decreases by 41% at the third layer comparedto the single layer. Figure 3d shows the relationship betweenNMSEs and thenumber of reservoir states used to generate output in final layer (NOUT) forthe two networks. Network 2 (deep and parallel layered) clearly shows areduction in errors compared to Network 1(deep layered). This is possiblydue to the fact that the numberof nodeswas effectively increased,while suchincrease is generally difficult to achieve with physical reservoirs, as men-tioned above. However, in Network 2 (deep and parallel layered), the errorincreased slightly at the fourth layer, which increase effectively halted theperformance improvement.Node Selection in deep-IGRThe decrease in performance with respect to the increase in the number ofnodes, which is shown in Fig. 3d, is thought to originate from the effect ofoverfitting due to the increase in the number of unnecessary nodes43.Overfitting in linear regression is determined by the combination ofreservoir size and training data length. Therefore, if the training data lengthis limited, overfitting will occur as the reservoir size increases (especially thenumber of unnecessary nodes). To verify this hypothesis, we analyzed thehigh dimensionality in the correlation coefficient rAB between node A andnode B shown in Eq. 4 for the reservoir state used at the output of the fourthlayer of this Network 224.rAB ¼PMk¼1 XA kð Þ � �XA� �XB kð Þ � �XB� �ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiPMk¼1 XA kð Þ � �XA� �2PMk¼1 XB kð Þ � �XB� �2q ð4Þwhere XA(k) is the reservoir state of node A and �XA is the average value ofXA(k). Figure 4a shows the correlation coefficients |rAB| between all reservoirstatesX Lð Þi (i = 1 ~ 200, L = 1 ~ 4) used to generate output in the fourth layerof Network 2 (i.e., 800 nodes in total). In the first layer, a randomwave u(k)was input to the IGR, whereas in layer L (≧2), the output y(L−1)(k) of theFig. 3 | Performance evaluation of a deep-iongating reservoir (deep-IGR) for Network 1 andNetwork 2 by the second-order nonlinear auto-regressive moving average (NARMA2) task.Schematic diagram of deep-IGR for (a) Network128–30 and (b)Network 228. cNMSEs of theNARMA2task vs. the number of layers of deep-IGR. The blackand blue plots show the results for Network 1 andNetwork 2, respectively. d Normalized meansquared errors (NMSEs) of theNARMA2 task vs thenumber of reservoir states used to generate output ineach layer.https://doi.org/10.1038/s44172-024-00227-y ArticleCommunications Engineering |            (2024) 3:81 5Fig. 4 | Node selection in deep-ion gating reservoir (deep-IGR). aHeatmap of thecorrelation coefficients of the reservoir states used in the output for the 4th layer ofNetwork 2. An example of (b) a strongly correlated or (c) relatively weakly correlatedreservoir state and a scatterplot. d The �rj������ plots, rearranging the node numbers j inorder from the lowest �rj������. e The relationship between the number of selected nodesN’ and the normalized mean squared error (NMSE) of the second-order nonlinearautoregressive moving average (NARMA2) task. f NOUT vs NMSEs ofNARMA2 task.https://doi.org/10.1038/s44172-024-00227-y ArticleCommunications Engineering |            (2024) 3:81 6previous layer was input to the IGR. Therefore, because these correlationcoefficients are small, due to the difference in inputs, the |rAB| in the firstlayer is very different from that in the rest of the layers, as shown in Fig. 4a.The inset of Fig. 4a shows an expanded part of the reservoir state in the firstlayer. The other major correlated/uncorrelated node combinations werealso achieved by the following node combinations. (1) gate current/draincurrent, (2) original pulse/inverted pulse, (3) virtual node corresponding topulse on time/pulse interval, etc. For example, the correlation coefficient |r|betweenX L¼1ð Þi¼5 andX L¼1ð Þi¼3 is close to 1, and they exhibit similar behavior, asshown in Fig. 4b. They correspond to different virtual nodes taken from thesame drain current (Lch = 5 µm).On the other hand,X L¼1ð Þi¼5 andX L¼1ð Þi¼11 havea relatively low correlation coefficient |r| of 0.48 and exhibit differentbehavior, as shown in Fig. 4c. These nodes correspond to different virtualnodes taken from different drain currents (Lch = 5 µm and Lch = 10 µm).To identify which of the 800 nodes shown in Fig. 4a has the lowercorrelation coefficient (i.e., contributes to the higher dimensionality), theaverage value of the correlation coefficient for node j (=1, …, 800) wascalculated as shown in Eq. 5;�rj������ ¼PNalli≠j rji������Nall � 1ð5ÞwhereNall is the number of all nodes (in this case,Nall = 800). The j�rjj plots,rearranging the node numbers j in order from the lowest j�rjj, are shown inFig. 4d. The lowest j�r1j is 0.3, whereas j�rjj increases with increasing j,indicating that j�rjj saturates at about j = 300. This suggests that about 500 ofthe total 800 node reservoir states are nodes that do not contribute to highdimensionality (i.e., the nodes are less effective for performing the giventask). To evaluate the effect of node correlation coefficients and highdimensionality on computational performance, the NARMA2 task wasperformed by increasing number of selected reservoir statesN’ in the orderof lower j�rjj (for example, when N’ = 100, Xj(j = 1 ~ 100) was used for thecalculation). Note thatN’ =NOUT in this case, where node selection is madeonly on the final layer (L = 4). Figure 4e shows the relationship between N’andNMSE; in the regionwhereN‘ is approximately 260 or less,N’ increaseswhile NMSE decreases for both training and testing. On the other hand, inthe region where N‘ is above approximately 260, N’ increases and thetraining error continues to decrease, while the test error increases. Also, asshown in Supplementary Fig. S7, the derivative of the test error converges toa slightly positive value in the region where N’ is above about 270. Thisindicates the effect of overfitting with increasing reservoir size, in whichreservoir states in regions of saturated diversity have a negative impact onthe computation. Figure 4f shows the results obtained with size reduction(Fig. 4e) and the relationship between theNMSEsof theNARMA2 task (testphase) and NOUT for Network 2. The red plots show the results of thecomputation when gradually excluding nodes, in order of largest j�rjj, fromthe 800 nodes which were used to generate output in the fourth layer ofNetwork 2. It was found that the NMSE decreased despite any reduction inthe number of reservoir states used to generate the output. Reducing toN’ = 600 (i.e., excluding from the calculation the 200 nodes with large j�rjj),the improvement in computational performancewas only about the sameasin the third layer without size reduction, but atN’ = 460 (i.e., excluding fromthe calculation the 340 nodes with large j�rjj), it rather outperformed aninterpolated version ofNetwork 2 without size reduction between L = 2 andL = 3. The performance continued to improve as the number of nodes wasreduced to about N’ = 300, but when N’ was further reduced, the perfor-mance rapidly degraded due to the loss of the nodes necessary for thecomputation being performed. These results indicate that the adverse effectsof multilayering, such as unnecessary node growth and overfitting, weresuccessfully suppressed without sacrificing the advantages of improvedcomputational performance due to multilayering. Here, network pruningwas performed by selectively incorporating uncorrelated nodes in the net-work through correlation coefficient analysis. This pruning methodeffectively exploits the feature of RC that its computational performance isachieved by high dimensionality. Therefore, the pruning method based onthe correlation coefficient is useful for identifying nodes that lead to higherdimension and identifying performance improvement mechanisms thatlead to higher dimension and performance improvement, although othermethods, such as network pruning based on network structure analysis canalso be used in our Deep-RC.Deep-IGR architecture, which utilizes node selection ateach layerThe node selection (reduction) by j�rjj was adapted to only one layer in theconfiguration shown in Fig. 4. Here, in order to maximize computationalperformance, we consider Network 3, in which node selection is adapted toall layers from the second layer onward by modifying Network 2 (deep andparallel layered). The number of selected reservoir states N’ used for theoutput of each layer was set to 400. Therefore, the first (single) and secondlayers, where the number of reservoir states NOUT is 200 and 400, respec-tively, do not perform node selection and are therefore identical to Network2, represented in Fig. 3b. The calculation method used for the third andsubsequent layers of Network 3 is explained below. The output in the thirdlayer of Network 2 used 600 reservoir states of X(L=1), X(L=2) and X(L=3) [here,these reservoir states are rewritten as Xj (j = 1, 2,…, 600)]. In Network 3, ofthese 600 reservoir states, N’ nodes (here 400 nodes) selected as X’(L=3) inorder of smallest j�rjj are used to generate the output of the third layer y(L=3).Furthermore, by voltage converting y(L=3) and inputting it to IGR again,X(L=4)(200 nodes) is obtained, and together with the reservoir states X’(L=3) (400nodes) obtained in the previous layers, there are a total of 600 reservoir states[here, these reservoir states are rewritten asXj (j = 1, 2,…, 600)]. As above,N’(=400) of these 600 nodes with small j�rjj were selected X’(L=4) and used togenerate output in the fourth layer y(L=4). Figure 5a and b show schematicdiagrams of how the reservoir outputs y(L) of Network 2 and Network 3 inlayer L are computed, respectively. In Network 2, shown in Fig. 5a, thenumber of reservoir states used for output increases by 200 as the number oflayers increases, because all reservoir statesX(L=1),X(L=2),…,X(L) in layers 1 toL are used to generate the reservoir output in layer L. On the other hand, inNetwork3, shown inFig. 5b,X’(L) selectedbyN’nodes inorder of smallest j�rjjof the reservoir stateX’(L−1) selected up to the previous layer and the reservoirstate X(L) in Layer L, as described above. Therefore, the number of reservoirstates used to generate output isfixedatN’, regardless of thenumber of layersL (>1). Figure 5c shows a plot of NMSE vs. the number of layers in theNARMA2 task for Network 3 (deep and parallel layered with node selec-tion), with the number of selected nodesN’ = 400. The NMSE continued todecrease as the number of layers increased, for both training and testingerrors, with the fourth layer achieving better computational performancethan Network 2 (deep and parallel layered), shown by the dotted blue line,with an NMSE of 0.00326 in the training phase and 0.00921 in the testingphase. These values are 79% and 53% lower in the training and testingphases, respectively, compared to a single-IGR. Figure 5d shows the rela-tionship between the number of nodes used for output NOUT and NMSE(test phase). Network 3 showed improved computational performance overthe conventional Network 2, despiteNOUT after the second layer being fixedat 400 ( =N‘). This performance improvement is explained by the networkbeing trained by increasing the number of layers while excluding unneces-sary nodes that cause overfitting and nodes that do not contribute todiversity, thereby effectively incorporating the features obtained in eachlayer. The predicted and target waveforms at layers 1 and 4 are shown inFig. 5e and f, respectively. Although the predictedwaveform in thefirst layercaptured the trend of the target waveform, there was a divergence betweenthepredictedwaveformand the targetwaveform in someareas.On theotherhand, the predicted waveforms at the fourth layer, shown in Fig. 5f, are inalmost perfect agreement. This shows that the deep-IGRbasedonNetwork3is able to successfully solve theNARMA2model shown inEq. 2, and that thisarchitecture drastically improved the computing performance of the IGR.Supplementary Fig. S8 shows the N’ dependence of NMSE during the testphase of the 3-layerDeep-IGR (Network3). For thedeep layered scheme, thehttps://doi.org/10.1038/s44172-024-00227-y ArticleCommunications Engineering |            (2024) 3:81 7test error was minimal when N’ = 400, therefore, N’ = 400 is set here. Fig-ure 6a shows a comparison of the computational performance evaluated inthe NARMA2 task of this study with other physical reservoirs reported sofar22–25,38–40.Here, inorder tomaximize theperformanceof thedeep-IGR, theoutputweights of thefinal layer (L = 4) ofNetwork 3were trainedwith ridgeregression. Our deep-IGR (Network 3, L = 4) achieved NMSE = 9.08×10−3in the test phase for the NARMA2 task (NMSE = 3.37 × 10−3 in the trainingphase), andachieved thehighest computational performanceof anyphysicalreservoir reported to date for the NARMA2 task22–25,38–40, even out-performing the NMSE = 0.016 of fully simulated reservoir computing[Standard Echo State Network (ESN)]38. This standard ESN is a single-reservoir with 300 nodes of a hyperbolic tangent as the activation function38,and differs from Deep-IGR (400 nodes and four multilayers) in its networkarchitecture and reservoir size. Although they are not a perfect like-for-likecomparison, ESN, which is a typical example of a simulation reservoir, hasbeenwidely compared in performancewithmany physical reservoirs1,38,44–46.Simulation reservoirs that canadjust thenetwork structure and reservoir sizeaccording to the purpose generally perform better than PRCs, and actuallyno PRCs that outperform the reported standard ESNs have been reportedfor theNARMA2 task, as shown in Fig. 6a. Therefore, this comparison has acertain value because our Deep-IGR is the first physical system to outper-form this reported simulation reservoir in the NARMA2 task22–25,38–40. Inparticular, the NARMA2 task is one of the best tasks for evaluating andcomparing the basic information processing capabilities of PRC, since thetask of analyzing second-order dynamic systems is widely performed byPRC1,10,13,22–25,38–40,47,48. On the other hand, the NARMA10 task is consideredas a more challenging benchmark task, but Deep-IGR was unable to solvethe NARMA10 task. This is due to the relatively small memory capacity(MC) of IGR (MC~ 4). It is expected that if our Deep-RC architecture isimplemented in adynamical systemwithhighmemory capacity (e.g., opticalelements and spin-wave interference)5,6,8,25, the NARMA10 task can besolvedwithhigh accuracy.Wewould like to demonstrate this as futurework.The deep-RC architecture proposed in this study has very few hyper-parameters, and can easily enhance the computational performance ofphysical reservoirs. Usually, it is extremely difficult to maximize the infor-mation processing capability of a physical system in physical reservoircomputing. There is a wide variety of parameters to be considered if the bestcomputational performance is to be obtained from a physical reservoir;Fig. 5 | Performance evaluation of a deep-iongating reservoir (deep-IGR) for Network 3 by thesecond-order nonlinear autoregressive movingaverage (NARMA2) task. Schematic diagram oflayer L of a deep-IGR for (a) Network 2 and (b)Network 3 with node selection. c Normalized meansquared errors (NMSEs) of the NARMA2 task vs.the number of layers of deep-IGR. d NMSEs of theNARMA2 task vs NOUT. Target waveform andpredicted waveform by IGR at (e) L = 1 and (f) L = 4.https://doi.org/10.1038/s44172-024-00227-y ArticleCommunications Engineering |            (2024) 3:81 8these include the delay time of the feedback loop, the mask matrix used forinput signal preprocessing (mask format, number of masks, mask length,etc.), in addition to the intensity and frequency of the signal input to thephysical system. Such complexity makes it almost impossible to experi-mentally search for the best combination of driving conditions for a physicalsystem.Hence, the deep-layered physical reservoir architecture proposed inthis study can dramatically improve the computational performance, andcan be realized in a small parameter space that is possible to realisticallyexplore. We experimentally examined three different deep reservoir archi-tectures, which are shown in Fig. 6b-d, and found that Network 1 (deeplayered, shown in Fig. 6b), which was previously reported for simulatedreservoirs, does not necessarily contribute to improved performance inphysical reservoirs28–30. Said scheme is suitable only as a component part ofphysical reservoirs, whose characteristics are not so sensitive to operatingconditions. Further, it does not increase the network size, which is a seriousproblem faced by physical reservoirs. On the other hand,we have succeededin increasing the network size and in improving the computational per-formance of the physical reservoir drastically with Network 3 (deep andparallel layeredwithnode selection),which is amodified versionofNetwork2 (deep and parallel layered, shown in Fig. 6c) that selectively uses usefulnodes for information processing, as shown in Fig. 6d. These deep-RCschemes should be practically implemented by integrating an FPGA andIGRs, as shown in Supplementary Fig. S9, whichwewould like to address asfuture work. Our deep-IGR is the first nanodevice implementation of deep-RC, and the dramatic improvement of the performance of IGR by ourarchitecture provides the possibility of realizing large-scale, brain-likephysical devices.ConclusionIn this study, we demonstrated the implementation of Deep-RC by nano-devices and the improvement of computational performance achieved byoptimizing the deep network structure, whereas Deep-RC has only beenimplemented by simulated RCs and limited physical reservoirs. Weimplemented IGR that uses ion-electron coupling dynamics as a compu-tational resource23 in the reservoir part, so as to create a multi-layeredframework, and found that a simple serial network structure did notimprove the performance (Network 1). On the other hand, we determinedthat a network structure that uses the reservoir state of the previous layer foroutputs shows a remarkable improvement in performance (Network 2).This is because, in addition to succeeding in further increasing the dimen-sion, which is generally difficult in physical systems, the predicted output ofthe reservoir is again input to the IGR and transferred to the high-dimensional feature space, where it is learned so that the deviation from thetarget output is reduced. In order to suppress overlearning due to increase ofunnecessary nodes, which was confirmed in this network structure, nodeselection based on correlation coefficients was used in Network 3 to effec-tively extract nodes that are effective for information processing. As a result,an NMSE of 9.08 × 10−3 was achieved in the NARMA2 task for deep-IGR.This is the best performance of all the physical reservoirs reported so far fortheNARMA2prediction taskwhich is a typical benchmark task forRCs andPRCs, and even outperforms full-simulation RC22–25,38–40. The easyand dramatic performance improvement of IGRs with our deep-RCarchitecture, and opens the way to the implementation of practical, large-scale, brain-based physical devices.Fig. 6 | Performance comparison with other phy-sical reservoirs and summary of deep-reservoircomputing (deep-RC) structures in this study.a Performance comparison with other physicalreservoirs by the second-order nonlinear auto-regressive moving average (NARMA2) task22–25,38–40.The result for soft body22 was converted to thenormalizedmean squared error (NMSE) in Eq. 3 forcomparison. Schematic diagram of the Deep-RCarchitecture for physical reservoir computing (PRC)investigated in this study for (b) Network 128–30, (c)Network 228 and (d) Network 3.https://doi.org/10.1038/s44172-024-00227-y ArticleCommunications Engineering |            (2024) 3:81 9MethodDevice fabrication and electrical measurementsHydrogen-terminated diamond homoepitaxial film was deposited, as achannel, on a IIa-type high-pressure high-temperature single-crystal dia-mond substrate (100) (EDP) by the microwave-plasma chemical vapordeposition (MPCVD) method. During the deposition process, 500 and 0.5standard cubic centimeters per minute of H2 and CH4, respectively, weresupplied, and the hydrogen-terminated diamond was grown at a radiofrequency of 950W. The IGR were fabricated with nine different channellengths (5, 10, 25, 35, 50, 100, 250, 350, and500μm), allwith a channelwidthof 50 μm. Pd/Pt electrodes (10 and 35 nm, respectively) were deposited assource and drain electrodes by electron beam evaporation with masklesslithography after oxygen termination of the diamond surfaces (other thanchannels) by oxygen plasma asher. A 3.5-μm LSZO thin film, used as anelectrolyte, was deposited by pulsed laser deposition (PLD) with an ArFexcimer laser. A 100-nm Au thin film was deposited by electron beamdeposition as a gate electrode.Electricalmeasurements of IGRwereperformedby the sourcemeasureunit and pulse measure unit of a semiconductor parameter analyzer(4200A-SCS, Keithley), which measurements were carried out at roomtemperature inside a vacuum chamber that had been evacuated by a turbomolecular pump. Probers were used to connect the IGR inside the chamber.Linear regression algorithm for learning readout networksIn theNARMA2 task,whichwas used in this study to evaluate performance,the readout network was trained with linear regression using the algorithmdescribed below. The reservoir output shown in Eq. 1 can also be expressedas;y kð Þ ¼ W � x kð Þ ð6ÞwhereW = (b,W1,W2,…,WN) and x(k) = [1, X1(k), X2(k),…, XN(k)]T arethe weight vector and the reservoir state vector, respectively. The reservoiroutput Y for all training periods (k = 1, 2,…,M) is described by;Y ¼ WX ð7ÞwhereX = (x(1), x(2),…, x(M)) andM = 1600 are the reservoir statematrixand the data length for the training phase, respectively. The weight matrixthat minimizes the squared error is given by;W ¼ YtXy ð8Þwhere Xy ¼ ½XT XXT� ��1� and Yt are the Moore-Penrose pseudo-inversematrix and the targetmatrix, respectively. In the results shown in Fig. 6a, thereadout weights of the final layer of Deep-IGR (Network 3) were trainedwith the ridge regression as follows.W ¼ YtXT XXT þ λI� ��1 ð9Þwhere λ (=2 × 10−3) and I ð� RN ×N Þ are the ridge parameters and identitymatrix, respectively. The learning and the sum-of-product calculationperformed in the readout network were done on a personal computer usingcurrent data obtained from the IGR.However, by utilizing artificial synapticdevices that reproduce weights by conductance, the sum-of-productcalculation performed in the readout can also be calculated in a physicalprocess, which is expected to further improve efficiency49–58.Data availabilityThe datasets generated during and/or analyzed during the current study areavailable from the corresponding author on reasonable request.Code availabilityThe codes used in the current study are available from the correspondingauthor on reasonable request.Received: 6 October 2023; Accepted: 11 June 2024;References1. Tanaka, G. et al. Recent advances in physical reservoir computing: Areview. Neural Netw. 115, 100–123 (2019).2. Nakajima, K. Physical reservoir computing—an introductoryperspective. Jpn. J. Appl. Phys. 59, 060501 (2020).3. Jaeger, H. The ‘echo state’ approach to analysing and trainingrecurrent neural networks-with an Erratum note 1.Germany.Ger. NatlRes. Cent. Inf. Technol. GMD Tech. Rep. 148, 13 (2001).4. Jaeger, H. & Haas, H. Harnessing Nonlinearity: Predicting ChaoticSystems and Saving Energy in Wireless Communication. Science304, 78–80 (2004).5. Paquot, Y. et al. Optoelectronic reservoir computing. Sci. Rep. 2,287 (2012).6. Van der Sande, G., Brunner, D. & Soriano,M. C. Advances in photonicreservoir computing. Nanophotonics 6, 561–576 (2017).7. Torrejon, J. et al. Neuromorphic computing with nanoscale spintronicoscillators. Nature 547, 428–431 (2017).8. Nakane, R., Tanaka, G. & Hirose, A. Reservoir computing withspin waves excited in a garnet film. IEEE A ccess 6, 4462–4469(2018).9. Tsunegi, S. et al. Physical reservoir computing based on spin torqueoscillator with forced synchronization. Appl. Phys. Lett. 114,164101 (2019).10. Jiang,W. et al. Physical reservoir computing usingmagnetic skyrmionmemristor and spin torque nano-oscillator. Appl. Phys. Lett. 115,192403 (2019).11. Akashi, N. et al. Input-driven bifurcations and informationprocessing capacity in spintronics reservoirs. Phys. Rev. Res. 2,043303 (2020).12. Sillin, H. O. et al. A theoretical and experimental study ofneuromorphic atomic switch networks for reservoir computing.Nanotechnology 24, 384004 (2013).13. Du, C. et al. Reservoir computing using dynamic memristors fortemporal information processing. Nat. Commun. 8, 2204 (2017).14. Moon, J. et al. Temporal data classification and forecasting using amemristor-based reservoir computing system. Nat. Electron. 2,480–487 (2019).15. Midya, R. et al. Reservoir computing using diffusive memristors. Adv.Intell. Syst. 1, 1900084 (2019).16. Zhu, X., Wang, Q. & Lu,W. D. Memristor networks for real-time neuralactivity analysis. Nat. Commun. 11, 2439 (2020).17. Sun, L. et al. In-sensor reservoir computing for language learning viatwo-dimensional memristors. Sci. Adv. 7, eabg1455 (2021).18. Hochstetter, J. et al. Avalanches and edge-of-chaos learning inneuromorphic nanowire networks. Nat. Commun. 12, 4008 (2021).19. Zhong, Y. et al. Dynamic memristor-based reservoir computing forhigh-efficiency temporal signal processing. Nat. Commun. 12,408 (2021).20. Milano, G. et al. In materia reservoir computing with a fully memristivearchitecture based on self-organizing nanowire networks.Nat. Mater.21, 195–202 (2022).21. Nakayama, J., Kanno, K. & Uchida, A. Laser dynamical reservoircomputing with consistency: an approach of a chaos mask signal.Opt. Express 24, 8679–8692 (2016).22. Nakajima, K., Hauser, H., Li, T. & Pfeifer, R. Information processing viaphysical soft body. Sci. Rep. 5, 10487 (2015).23. Nishioka, D. et al. Edge-of-chaos learning achieved by ion-electron–coupled dynamics in an ion-gating reservoir. Sci. Adv. 8,eade1156 (2022).24. Wada, T. et al. A Redox-based Ion-Gating Reservoir, Utilizing DoubleReservoir States in Drain and Gate Nonlinear Responses. Adv. Intell.Syst. 5, 2300123 (2023).https://doi.org/10.1038/s44172-024-00227-y ArticleCommunications Engineering |            (2024) 3:81 1025. Namiki, W. et al. Experimental Demonstration of High-PerformancePhysical Reservoir Computing with Nonlinear Interfered Spin WaveMulti-Detection. Adv. Intell. Syst. 5, 2300228 (2023).26. Takayanagi, M. et al. Ultrafast-switching of an all-solid-state electricdouble layer transistor with a porous yttria-stabilized zirconia protonconductor and the application to neuromorphic computing.Mater.Today Adv. 18, 100393 (2023).27. Appeltant, L. et al. Information processing using a single dynamicalnode as complex system. Nat. Commun. 2, 468 (2011).28. Hasegawa, H., Kanno, K. & Uchida, A. Parallel and deep reservoircomputing using semiconductor lasers with optical feedback.Nanophotonics 12, 869–881 (2023).29. Akiyama, T. & Tanaka, G. Computational Efficiency of Multi-StepLearning Echo State Networks for Nonlinear Time Series Prediction.IEEE Access 10, 28535–28544 (2022).30. Akiyama, T. & Tanaka, G. Analysis on Characteristics of Multi-StepLearning Echo State Networks for Nonlinear Time Series Prediction;Analysis on Characteristics of Multi-Step Learning Echo StateNetworks for Nonlinear Time Series Prediction. 2019 InternationalJoint Conference on Neural Networks (IJCNN) 1–8 (2019).31. Gallicchio, C. & Micheli, A. Echo State Property of Deep ReservoirComputing Networks. Cogn. Comput. 9, 337–350 (2017).32. Goldmann, M., Köster, F., Lüdge, K. & Yanchuk, S. Deep time-delayreservoir computing: Dynamics and memory capacity. Chaos 30,093124 (2020).33. Nakajima, M. et al. Physical deep learning with biologically inspiredtraining method: gradient-free approach for physical hardware. Nat.Commun. 13, 7847 (2022).34. Liu, K. et al. Multilayer Reservoir ComputingBasedon Ferroelectric α-In2Se3 for Hierarchical Information Processing. Adv. Mater. 34,2108826 (2022).35. Lin, B.-D. et al. Deep time-delay reservoir computing with cascadinginjection-locked lasers. IEEE J. Sel. Top. Quantum Electron. 29,1–8 (2022).36. Wright, L. G. et al. Deep physical neural networks trained withbackpropagation. Nature 601, 549–555 (2022).37. Tsuchiya, T. et al. The electric double layer effect and its strongsuppression at Li+ solid electrolyte/hydrogenated diamondinterfaces. Commun. Chem. 4, 117 (2021).38. Kan, S. et al. Simple reservoir computing capitalizing on the nonlinearresponse of materials: theory and physical implementations. Phys.Rev. Appl. 15, 024030 (2021).39. Kan, S., Nakajima, K., Asai, T. & Akai‐Kasaya, M. Physicalimplementation of reservoir computing through electrochemicalreaction. Adv. Sci. 9, 2104076 (2022).40. Akai-Kasaya, M. et al. Performance of reservoir computing in arandom network of single-walled carbon nanotubes complexed withpolyoxometalate. Neuro. Comput. Eng. 2, 014003 (2022).41. Inubushi, M. & Yoshimura, K. Reservoir computing beyond memory-nonlinearity trade-off. Sci. Rep. 7, 10199 (2017).42. Atiya, A. F. & Parlos, A. G. New results on recurrent network training:unifying the algorithms and accelerating convergence. IEEE Trans.Neural Netw. 11, 697–709 (2000).43. Dambre, J., Verstraeten, D., Schrauwen, B. & Massar, S. Informationprocessing capacity of dynamical systems. Sci. Rep. 2, 514 (2012).44. Maksymov, I. S. Analogue and physical reservoir computing usingwaterwaves:Applications inpowerengineeringandbeyond.Energies16, 5366 (2023).45. Maksymov, I. S. & Pototsky, A. Reservoir computing based onsolitary-like waves dynamics of liquid film flows: A proof of concept.Europhys. Lett. 142, 43001 (2023).46. Maksymov, I. S., Pototsky, A. & Suslov, S. A. Neural echo statenetwork using oscillations of gas bubbles in water. Phys. Rev. E 105,044206 (2022).47. Nishioka, D., Shingaya, Y., Tsuchiya, T., Higuchi, T. & Terabe, K. Few-and single-molecule reservoir computing experimentallydemonstrated with surface enhanced Raman scattering and ion-gating. Sci. Adv. 10, eadk6438 (2024).48. Shibata, K. et al. Redox-based ion-gating reservoir consisting of (104)oriented LiCoO2 film, assisted by physical masking. Sci. Rep. 13,21060 (2023).49. Arnold, A. J. et al. Mimicking neurotransmitter release in chemicalsynapses via hysteresis engineering in MoS2 transistors. ACS Nano11, 3110–3118 (2017).50. Ielmini, D. & Wong, H. S. P. In-memory computing with resistiveswitching devices. Nat. Electron. 1, 333–343 (2018).51. Ielmini, D. Brain-inspired computing with resistive switching memory(RRAM): Devices, synapses and neural networks.Microelectron. Eng.190, 44–53 (2018).52. Zhu, J. et al. Ion gated synaptic transistors basedon 2DvanderWaalscrystals with tunable diffusive dynamics. Adv. Mater. 30, 1800195(2018).53. Kumar,S.,Williams,R.S.&Wang,Z. Third-order nanocircuit elementsfor neuromorphic engineering. Nature 585, 518–523 (2020).54. Schranghamer, T. F., Oberoi, A. & Das, S. Graphene memristivesynapses for high precision neuromorphic computing.Nat. Commun.11, 5474 (2020).55. Sebastian, A. et al. Two-dimensional materials-based probabilisticsynapses and reconfigurable neurons for measuring inferenceuncertainty using Bayesian neural networks. Nat. Commun. 13,6139 (2022).56. Wu, X., Dang, B., Wang, H., Wu, X. & Yang, Y. Spike‐Enabled AudioLearning inMultilevel SynapticMemristor Array‐BasedSpikingNeuralNetwork. Adv. Intell. Syst. 4, 2100151 (2022).57. Kumar, S., Wang, X., Strachan, J. P., Yang, Y. & Lu, W. D. Dynamicalmemristors for higher-complexity neuromorphic computing.Nat.Rev.Mater. 7, 575–591 (2022).58. Nishioka, D., Tsuchiya, T., Higuchi, T. & Terabe, K. Enhanced synapticcharacteristics of HxWO3-based neuromorphic devices, achieved bycurrent pulse control, for artificial neural networks. Neuromorph.Comput. Eng. 3, 034008 (2023).AcknowledgementsThis work was in part supported by Japan Society for the Promotion ofScience (JSPS) KAKENHI Grant Number JP22H04625 (Grant-in-Aid forScientific Research on Innovative Areas “Interface Ionics”), andJP22KJ2799 (Grant-in-Aid for JSPS Fellows). A part of this work wassupportedbyJSTPRESTOGrantnumber, JPMJPR23H4.Apart of thisworkwas supported by “Advanced Research Infrastructure for Materials andNanotechnology in Japan (ARIM)” of the Ministry of Education, Culture,Sports, Science and Technology (MEXT). Proposal NumberJPMXP1223NM5072. A part of this work was supported by the IketaniScience and Technology Foundation.Author contributionsD.N., T.T., and K.T. conceived the idea for the study. D.N. and T.T. designedthe experiments. D.N. and T.T. wrote the paper. D.N. carried out theexperiments.D.N. prepared the samples.D.N. andT.T. analyzed thedata.Allauthors discussed the results and commented on the manuscript. K.T.directed the projects.Competing interestsThe authors declare no competing interests.Additional informationSupplementary information The online version containssupplementary material available athttps://doi.org/10.1038/s44172-024-00227-y.https://doi.org/10.1038/s44172-024-00227-y ArticleCommunications Engineering |            (2024) 3:81 11https://doi.org/10.1038/s44172-024-00227-yCorrespondence and requests for materials should be addressed toTakashi Tsuchiya.Peer review informationCommunications Engineering thanks Flavio AbreuAraujo, Anatole Moureaux, Fabien Alibart, and the other, anonymous,reviewer(s) for their contribution to the peer review of this work. PrimaryHandling Editors: Damien Querlioz, Anastasiia Vasylchenkova, RosamundDaw.Reprints and permissions information is available athttp://www.nature.com/reprintsPublisher’snoteSpringerNature remainsneutralwith regard to jurisdictionalclaims in published maps and institutional affiliations.Open Access This article is licensed under a Creative CommonsAttribution 4.0 International License, which permits use, sharing,adaptation, distribution and reproduction in anymedium or format, as longas you give appropriate credit to the original author(s) and the source,provide a link to the Creative Commons licence, and indicate if changeswere made. The images or other third party material in this article areincluded in the article’s Creative Commons licence, unless indicatedotherwise in a credit line to the material. If material is not included in thearticle’sCreativeCommons licence and your intended use is not permittedby statutory regulation or exceeds the permitted use, you will need toobtain permission directly from the copyright holder. To view a copy of thislicence, visit http://creativecommons.org/licenses/by/4.0/.© The Author(s) 2024https://doi.org/10.1038/s44172-024-00227-y ArticleCommunications Engineering |            (2024) 3:81 12http://www.nature.com/reprintshttp://creativecommons.org/licenses/by/4.0/ A high-performance deep reservoir computer experimentally demonstrated with ion-gating reservoirs Results Ion-gating reservoirs using electric double-layer transistors Performance evaluation of deep ion-gating reservoir with NARMA2�task Node Selection in deep-IGR Deep-IGR architecture, which utilizes node selection at each�layer Conclusion Method Device fabrication and electrical measurements Linear regression algorithm for learning readout networks Data availability Code availability References Acknowledgements Author contributions Competing interests Additional information