Agustin Sansone

Privat

Clear skies and a rainbow in front of Cologne Cathedral

“Receiving support from the DAAD-Stiftung made it possible for me to conduct research in Germany, expand my academic skills, and experience life in a new culture. Both professionally and personally, it was an incredibly enriching journey.”

As part of his doctoral studies, Agustín Sansone researched the reconstruction of the human brain using deep neural networks. With support of the UNICORE Scholarship by the DAAD-Stiftung, he was able to further develop his project at the Forschungs-zentrum Jülich, primarily due to its extensive data resources and advanced computing capabilities. He found Germany to be very hospitable and could imagine living here in the future.

The main goals of my research project was to develop deep neural networks to generate 3D point cloud brain configurations based on Laplace-Beltrami shape descriptors (LBSDs). These descriptors, derived from differential geometry, encapsulate geometric and topological properties intrinsic to 3D objects. Systematic exploration of alterations to LBO features could provide insights into brain geometry and its relation to individual-level characteristics. As our main hypothesis, we expect that the concepts presented by Marin et al. in 2021 can be applied to our context: the human brain.

I attended German schools throughout my education in Argentina, so I didn’t have to learn the language from scratch at an older age. When I arrived in Germany, I felt like I had lost much of my German, but after a few weeks, it started coming all back naturally. I believe that with some focused study and practice over a few months, I can surpass my previous C1 level.

Sansone FRA

Privat

Arriving at Frankfurt Airport

As a workplace, Forschungszentrum Jülich is pretty standard in terms of infrastructure: each person has an office with whiteboards, conference rooms, and everything needed for work. In our field, all we really need is a computer, so there was no need for a high-tech lab filled with equipment. Of course, there was at our disposal a lot of hardware that we could use to train our ML models. I particularly liked that the Jülich campus is composed of many buildings from various scientific disciplines.

About the project, we managed to adapt and apply concepts from previous work to our specific context of the human brain. Additionally, I had the opportunity to work with large computing clusters, which is something I simply cannot do in my lab in Buenos Aires. Our university lacks resources like GPUs and typically offers only cheap computers to PhD students. I gained quality experience training large models using many GPUs in parallel.

Sansone Jülich

Privat

At Forschungszentrum Jülich

Regarding my future path and the influence of the Forschungszentrum Jülich, I can say this research-stay definitely has influenced my academic path. I now have more practical tools and experience.

Additionally, I had the chance to see how research is done in another lab in a completely different region of the world, which broadened my perspective on how other labs do science. The research group in Jülich was a friendly group of very international people.

Sansone Rhein

Privat

Visiting Cologne, view at the river Rhine

In terms of personal life, most cities in Germany felt quieter and calmer compared to Buenos Aires, where I’m from. I loved seeing people enjoying parks, walks along the river, bars, and events at all hours.

I also felt extremely safe, unlike Argentina, where it is dangerous to walk alone at night. I also liked the WG (shared flat) culture. It allowed me to interact with Germans daily and learn a lot.

Sansone Wurst

Privat

Sausage is never far away: German Currywurst and a sausage vending machine

After this experience, I could see myself living in Germany long term. Salaries are higher, the culture is rich, and I found the locals to be kind and welcoming. It’s also much safer and stable for raising a family. And the education system impressed me: most Germans I met speak at least three languages and are curious and well-educated.

Sansone PoppSchloß

Privat

A trip to Bonn: Poppelsdorfer Schloss

I hope that some colleagues from my lab in Argentina can come to Jülich for an exchange in future collaborations. The resources are very valuable and not available at our home lab, and since we all work on similar topics, it would be beneficial for everyone.

As of July 2025.

Read in detail about the technical aspects of Agustín Sansone's research:

Data

We used data from the UK Biobank, which includes MRI scans and personal characteristics from around 50,000 individuals. After some filtering we considered 20,998 healthy subjects for training and 5,256 for testing.

Preprocessing

FreeSurfer was used for segmenting brain structures (e.g., hippocampus, amygdala, caudate etc.). Then BrainPrint produces 3D point clouds, fits a mesh and applies the Laplace-Beltrami operator (LBO) decomposition to obtain the LBSDs.

Once we have the computed meshes for each brain region, we convert them into an 86x86x86 pixel image. The preprocessed image contains a 1 or 0 on each pixel symbolizing whether that point in space is inside or outside that brain region represented by the mesh. Since each pixel matches with the granularity measured by the actual raw MRI scan, there is no information loss by this conversion.

Model Architecture

At the core of this approach lies an AutoEncoder (AE) architecture, where the 3D point cloud (X) and its corresponding Laplacian spectrum (Spec(X)) are inputted into the network. Through the AE, one aims to enforce a reconstruction such that X is approximately equal to X̃. Here, X represents the brain structure, Spec(X) denotes its spectrum.

The architecture consists of four key blocks: the encoder (E), decoder (D), spectral encoder (π), and translation (ρ), all of which are learnable and parametrized by the neural network. The model's objective is to enable the recovery of the brain input purely from its eigenvalues via the composition D(π(Spec(X))) ≈ X.

We studied different architectures for every main block. What tended to work the best was the integration of Convolutional Neural Networks (CNNs) into the architecture, which naturally addresses the loss of spatial information inherent in 3D structures when using fully connected networks.

Skip connections were also useful, since it helps with the vanishing gradient problem and allows to train deeper networks. The addition of dropout into the layers was key to prevent overfitting for this problem. Batch normalization layers were also very sensitive to this dataset, the absence of some of them made the model get easily stuck in local minimums.

Some feature engineering was tried by doing some calculations among consecutive eigenvalues, but that idea did not improve the predictions in any case. It was tested to forwardly predict the final image from the spectrum without some blocks like the Encoder, or ρ block, but empirically it was never better to remove these blocks from the architecture.

Final Hyperparameters

To be consistent with previous work we considered training with the first 50 eigenvalues, although the first 100 eigenvalues were computed. In our final setup, we trained the models for 20 epochs, starting with a learning rate of 0.1 and decreasing it to 0.01 after the 10th iteration. The batch size is 512 (a bigger one was not possible due to hardware limitations). The optimizer is Adam and the criterion is MSE. On each layer of the network the ReLu activation is used, with the exception of the last layer in the Decoder, where sigmoid is used.

The dropout factor on each convolutional layer is 0.1 and the momentum on each batch normalization layer is 0.01. The box as input and output has a size of 86, while the latent space box has a size of only 5. The number of internal layers on the π, ρ, Encoder and Decoder block is 6.

With the current final setup it takes between 2 and 3 days to train a model using 4 GPUs in parallel. The final model has 640,852 trainable parameters. The split for train/test was set to 80% and 20%.

Loss function

The loss function defined by Marin et al. is: |E(X)-π(Spec(X))| + c * |D(E(X))-X| + c * |ρ(E(X))-Spec(X)| Where c is a small constant value. We alternatively proposed this loss function family: |D(π(Spec(X)))-X| + c * |E(X)-π(Spec(X))| + c * |ρ(π(Spec(X)))-Spec(X)|

Notice the latent-space term is now not the most important, and we replaced |D(E(X))-X| by |D(π(Spec(X)))-X|, which is actually how the prediction in the end is computed. Similarly we also replaced ‘ρ(E(X))’ by ‘ρ(π(Spec(X)))’, but that did not seem to make a big impact on the results.

As one could expect, with this alternative the prediction error is lower, and the learned latent space error is higher.

We noticed that some regions were much bigger than others. This problem is conceptually similar to class imbalance in classification tasks. To mitigate it, we introduced a weighted loss function.

Pending ideas

One idea is, given a fully trained model, to manually change some specific eigenvalues on the input to investigate how it affects the prediction. There is some previous work that correlates the width, height and depth of the structures with the first three eigenvalues. So it would be interesting to see how one could experiment on these relationships by artificially playing with the eigenvalues.

The other discussed idea has to do with what is known as concept proof. This is a general approach on the field that helps give bit complex models some degree of interpretability. It involves probing the trained neural network to assess whether it has learned some human-understandable "concepts" or not.

This is done by fitting linear-like classifiers to the activations at various layers to see if a given concept (e.g. "the brain region is symmetrical", "it has a hole inside", “it is more inside or near to the cortex”, “it is more on the left or right hemisphere”, “is near a specific region”) can be predicted from them. If the concept can be linearly decoded from a layer’s activations, it suggests that the network has internally represented that concept at that stage.

Future

The plan to polish these partial findings, implement the pending ideas and afterwards send the work into a related journal.