{"id":8014,"date":"2025-04-22T09:00:00","date_gmt":"2025-04-22T07:00:00","guid":{"rendered":"https:\/\/blog.structuralia.com\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones"},"modified":"2026-03-27T13:46:00","modified_gmt":"2026-03-27T12:46:00","slug":"proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones","status":"publish","type":"post","link":"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones","title":{"rendered":"Master's Thesis Project: A Model for Detecting Fraud in Transactions"},"content":{"rendered":"<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_84 counter-hierarchy ez-toc-counter ez-toc-white ez-toc-container-direction\">\n<div class=\"ez-toc-title-container\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\"><\/p>\n<span class=\"ez-toc-title-toggle\"><a href=\"#\" class=\"ez-toc-pull-right ez-toc-btn ez-toc-btn-xs ez-toc-btn-default ez-toc-toggle\" aria-label=\"Toggle Table of Contents\"><span class=\"ez-toc-js-icon-con\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewbox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewbox=\"0 0 24 24\" version=\"1.2\" baseprofile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/span><\/a><\/span><\/div>\n<nav><ul class='ez-toc-list ez-toc-list-level-1' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#Creacion_de_modelo_para_la_deteccion_de_fraudes_en_transacciones_en_tarjetas_de_credito\" >Development of a Model for Detecting Fraud in Credit Card Transactions<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#1_Eleccion_de_Dataset\" >1. Choosing a Dataset<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#2_ETL_con_Trifacta\" >2. ETL with Trifacta<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#1_Analisis_EDA\" >1. EDA Analysis<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#2_Estudio_de_correlaciones\" >2. Correlation Study<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#4_Visualizacion\" >4. Visualization<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#5_Construccion_de_modelo_de_deteccion_de_fraudes\" >5. Building a Fraud Detection Model<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#1_Definicion_de_variables_de_entrada_y_salida\" >1. Definition of input and output variables.<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#2_Arbol_de_decision\" >2. Decision Tree<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#3_Random_Forest\" >3. Random Forest<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-11\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#4_K-nearest_neighbors\" >4. K-nearest neighbors<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-12\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#5_Naive_bayes\" >5. Naive Bayes<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-13\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#_6_Comparacion_de_los_modelos\" >&nbsp;6. Comparison of the Models<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-14\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#7_Clustering\" >7. Clustering<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-15\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#8_Conclusiones\" >8. Conclusions<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-16\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#RESENA_DEL_AUTOR\" >AUTHOR'S REVIEW:<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-17\" href=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\/#TESTIMONIO_DEL_AUTOR\" >AUTHOR'S STATEMENT:<\/a><\/li><\/ul><\/nav><\/div>\n\n<p class=\"wp-block-paragraph\">Find out how a student has developed an innovative model for<strong> detect transaction fraud and combat cyberattacks<\/strong> in this master's thesis.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Currently, with the <strong>advances in technology and Internet access<\/strong>, payment methods have evolved, and there are now different ways to make payments, whether in person or online. Just as payment methods have changed to adapt to the needs and preferences of customers and businesses, the <strong>Cybercriminals have also been evolving<\/strong>. Despite the various security measures adopted by financial institutions, the <strong>payment fraud<\/strong> It's a problem we face every day.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The <strong>financial institutions and payment service providers<\/strong> must evolve in tandem with the sector's needs, making use of emerging technology and the availability of data to <strong>implement security measures<\/strong> and establish <strong>measures for early fraud detection.<\/strong><\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Creacion_de_modelo_para_la_deteccion_de_fraudes_en_transacciones_en_tarjetas_de_credito\"><\/span><strong>Development of a Model for Detecting Fraud in Credit Card Transactions<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">This interesting master's thesis, prepared <a href=\"https:\/\/www.linkedin.com\/in\/magaly-fonseca-94309152\/\">Magaly Fonseca<\/a> A former student of the Master's in Big Data and Business Analytics at Structuralia takes a deeper look at this topic\u2014which is so important today\u2014by creating a \u00abModel for Detecting Fraud in Credit Card Transactions.\u00bb.&nbsp;<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"1_Eleccion_de_Dataset\"><\/span><strong>1. Choosing a Dataset<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">For the preparation of this master's thesis <strong>A dataset was selected,<\/strong> from those available on the page <a href=\"https:\/\/www.kaggle.com\/datasets\/dhanushnarayananr\/credit-card-fraud\">Kaggle<\/a>, which contains data that can be used to train a machine learning model capable of detecting cases of fraud in transactions. The chosen dataset is called<strong> card_transdata<\/strong> which contains information on credit card transactions.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The dataset is a file with a .csv extension and <strong>It consists of 8 columns.<\/strong> Described below:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Distance_from_home:<\/strong> the distance from home to the location where the transaction took place.<\/li>\n\n\n\n<li><strong>Distance_from_last_transaction:<\/strong> indicates the distance between the transaction point and the last recorded transaction\u2014that is, the previous transaction.<\/li>\n\n\n\n<li><strong>Radio_to_median_purchase_price:<\/strong> It is the ratio of the transaction amount to the customer's average purchase price.<\/li>\n\n\n\n<li><strong>Repeat_retailer:<\/strong> Enter 1 if the transaction was made at the same retailer, and 0 if not.<\/li>\n\n\n\n<li><strong>Used_chip:<\/strong> In this column, a 1 indicates that the credit card chip was used for the transaction, and a 0 indicates that it was not used.<\/li>\n\n\n\n<li><strong>Used_pin_number:<\/strong> In this field, enter 1 if the PIN was used in the transaction, and 0 otherwise.<\/li>\n\n\n\n<li><strong>Online_order:<\/strong> It is marked with a 1 if the transaction corresponds to an online order and with a 0 if it does not.<\/li>\n\n\n\n<li><strong>Fraud:<\/strong> Indicates whether the transaction was flagged as fraudulent or not, just like the previous columns, with a value of 1 for fraudulent transactions and 0 for non-fraudulent ones.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"2_ETL_con_Trifacta\"><\/span><strong>2. ETL with Trifacta<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">For the <strong>data extraction, processing, and loading (ETL) process<\/strong> the tool is used <a href=\"https:\/\/www.trifacta.com\/\">Trifacta<\/a>. To start the process, load the card_transdata.csv dataset into the tool and create the flow with the name <em>\u201ctransaction data\u201d<\/em>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The first step in the transformation process is to <strong>data profiling<\/strong>, in order to gain a thorough understanding of the dataset's content and verify the quality of the data, with the goal of determining the transformations that need to be performed in order to work with the data and obtain reliable results.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">As shown in the image above, the data in the 8 columns of the dataset is of good quality, since the indicator bar is green in all cases, which means that the data is correct, matches the cell format, and contains no invalid or empty values. It can also be seen that the dataset contains two types of data: integers and decimals. Therefore, we proceed to complete the ETL process by running the job in Trifacta to obtain the dataset and move on to the analytical phase.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>3. Statistical Study<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An important part of working with data is understanding it and performing <strong>an analytical study <\/strong>that will allow us to make the best decisions regarding their treatment; for this process, the analysis was conducted in<strong> Python.<\/strong><\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"1_Analisis_EDA\"><\/span>1. EDA Analysis<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<figure class=\"wp-block-image size-full\"><img fetchpriority=\"high\" decoding=\"async\" width=\"507\" height=\"365\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-2-1-1.png\" alt=\"\" class=\"wp-image-12284\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-2-1-1.png 507w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-2-1-1-300x216.png 300w\" sizes=\"(max-width: 507px) 100vw, 507px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Figure 2. Graphs of the categorical variables.<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">As shown in the graphs, it can be seen that the <strong>88.2% of transactions<\/strong> were made at the same retailer as the previous purchase. In the <strong>The chip was used for 35% transactions.<\/strong> The PIN was used for 10.1% of purchases, and 65.1% of transactions were online purchases. Finally, the <strong>8.7% of transactions are fraudulent.<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Box plots are created for the continuous variables to examine the distribution of the data for those variables:<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" width=\"635\" height=\"316\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-3-1.png\" alt=\"\" class=\"wp-image-12285\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-3-1.png 635w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-3-1-300x149.png 300w\" sizes=\"(max-width: 635px) 100vw, 635px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Figure 3. Box plot of the continuous variables.<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For all three variables, outliers can be observed that stand out from the rest of the data when considering the variable descriptions, namely: distance from home, distance from the last transaction, and ratio relative to the average purchase amount; Furthermore, given that the goal is to train a model that predicts whether a transaction is fraudulent or not, these outliers could indicate fraudulent activity; therefore, it is determined that they should be retained.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"2_Estudio_de_correlaciones\"><\/span>2. Correlation Study<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">An important part of the analysis is conducting a correlation study to determine whether there is a relationship between the different variables in the dataset. Since the dataset contains both continuous and categorical variables, we use Spearman\u2019s correlation method, which can be applied to both types of variables:<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" width=\"472\" height=\"497\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-4-1.png\" alt=\"\" class=\"wp-image-12287\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-4-1.png 472w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-4-1-285x300.png 285w\" sizes=\"(max-width: 472px) 100vw, 472px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Figure 4. Correlation of the variables<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The result shows a<strong> strong correlation<\/strong>n (0.6) between the variables \u201cdistance from home\u201d and \u201cwhether the purchase was made at the same retailer.\u201d Furthermore, we can see a breakdown of the variables that influence the variable of interest (fraud), where the ratio relative to the average purchase amount is the variable most strongly correlated with fraud (0.3)\u2014although the value is not particularly high\u2014 followed by the \u201conline purchase\u201d variable at 0.2 and \u201cdistance from home\u201d at 0.1. An important point emerges from this analysis: two variables show a negative correlation with fraud\u2014chip use and PIN use\u2014both with a value of -0.1. This could be interpreted to mean that the use of these two security methods is effective in preventing fraud in transactions.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">As the final step in the statistical analysis, hypothesis tests (ANOVA) are conducted to determine whether or not the variables influence our variable of interest, which is fraud. The results are shown below:<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"420\" height=\"252\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-5-1.png\" alt=\"\" class=\"wp-image-12288\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-5-1.png 420w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-5-1-300x180.png 300w\" sizes=\"(max-width: 420px) 100vw, 420px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Figure 5. ANOVA Results<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">As can be seen, the only variable with a value greater than 0.05 is whether the purchase was made at the same retailer; therefore, the null hypothesis that the variable is influential is rejected. The remaining variables are indeed influential.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"4_Visualizacion\"><\/span>4. Visualization<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Once the ETL process and the statistical analysis of the available data have been completed, the dataset is loaded into <a href=\"https:\/\/powerbi.microsoft.com\/en-au\/\">PowerBI<\/a> to create a dashboard that displays the available information:<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"681\" height=\"350\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-6-1.png\" alt=\"\" class=\"wp-image-12289\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-6-1.png 681w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-6-1-300x154.png 300w\" sizes=\"(max-width: 681px) 100vw, 681px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Figure 6. Loading the dataset into PowerBI.<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Once the dataset has been loaded, we proceed to perform some <strong>data transformations<\/strong>, just to make it easier to see, these transformations apply to the data type of columns containing Boolean values; they are converted from 0 and 1 to the \u201ctrue\/false\u201d type, and in the case of the<strong> Fraud Column<\/strong> The data type is changed to text, and the 0s are replaced with the word \u201cno\u201d and the 1s with \u201cyes.\u201d It should be noted that these transformations are performed solely to facilitate visualization; they do not modify the data source (the dataset).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Using the transformations described above, we proceed to create the dashboard with the information relevant to the case study:<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"577\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-7-1-1024x577.png\" alt=\"\" class=\"wp-image-12283\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-7-1-1024x577.png 1024w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-7-1-300x169.png 300w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-7-1-768x432.png 768w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Imagen-7-1.png 1307w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">On the dashboard, you can see how <strong>Key figure: the total number of cases of fraud detected<\/strong>, which corresponds to 87,403 cases; this represents the <strong>8,741 TP3T transactions<\/strong>, which is clearly visible in the pie chart.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In addition, a table is provided that compares the medians of the numerical variables for fraud cases versus non-fraud cases; this table shows a significant difference in the medians of the variable <em>\u201cratio relative to the average purchase price\u201d<\/em>\u201din the <strong>cases of fraud<\/strong> (5.07) compared to the <strong>non-fraudulent transactions<\/strong> (0.91); there is also a significant difference in the medians of the variable \u201c<em>\u201ddistance from home\"<\/em>, since in cases of fraud the median is much higher than in cases of non-fraud; whereas for the variable \u201ctime since the last transaction,\u201d the difference between the medians is not significant (0.2).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Finally, the following are shown: <strong>4 bar charts<\/strong>, one for each Boolean variable related to the \"fraud\" variable. It is easy to see from these that, in cases of fraud, most of the transactions were carried out on the same<strong> retail<\/strong>, it is also evident that the largest proportion of fraud cases occurred in<strong> online shopping.<\/strong> Furthermore, it is evident that security measures such as the use of a chip and a PIN are effective in preventing fraud, with the use of a PIN being the most effective.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"5_Construccion_de_modelo_de_deteccion_de_fraudes\"><\/span><strong>5. Building a Fraud Detection Model<br><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">With the goal of conducting a <strong>fraud detection model<\/strong>, which would be a model of <strong>classification learning.<\/strong> First, we determine the variables to be used. As established at the end of the statistical analysis, the variables that influence whether a transaction is fraudulent or not are all of them\u2014except for whether the purchase was made at the same retailer. Furthermore, this variable had a high correlation with the \u201cdistance from home\u201d variable, so only one of them should be used; this provides yet another reason to exclude the variable. <em>\u201crepeat_retailer\u201d<\/em>.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"1_Definicion_de_variables_de_entrada_y_salida\"><\/span>1. Definition of input and output variables.<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Based on the information above, the input and output variables are defined as the first step in building the model; the training and test datasets are also defined, with the test set comprising 30% of the data; in addition, the data is scaled.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Once defined, the prediction is made using various supervised learning methods.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"2_Arbol_de_decision\"><\/span>2. Decision Tree<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"653\" height=\"213\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/IaY7I5JNQZ-jh0SDnHsHLG-cUISiWtHudgeHZUbj3H9fAfAT6pzNdeqvqOzaziZkuy95tNDyVRPw0m3nvsOpGIt4V2EjsxuJDgGE.png\" alt=\"\" class=\"wp-image-12282\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/IaY7I5JNQZ-jh0SDnHsHLG-cUISiWtHudgeHZUbj3H9fAfAT6pzNdeqvqOzaziZkuy95tNDyVRPw0m3nvsOpGIt4V2EjsxuJDgGE.png 653w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/IaY7I5JNQZ-jh0SDnHsHLG-cUISiWtHudgeHZUbj3H9fAfAT6pzNdeqvqOzaziZkuy95tNDyVRPw0m3nvsOpGIt4V2EjsxuJDgGE-300x98.png 300w\" sizes=\"(max-width: 653px) 100vw, 653px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">A confusion matrix is generated to verify the quality of the model<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"546\" height=\"98\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/RfzQSZ0u_6nSruQZiSRaosGwq9Zz5bEnwkoCNdbGTT-IT0u6gQJaOhKnskYsQ3gK-BQR2iaIDGMh0JAPQ4xPqDxd-oLO9eG64Ekw.png\" alt=\"\" class=\"wp-image-12281\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/RfzQSZ0u_6nSruQZiSRaosGwq9Zz5bEnwkoCNdbGTT-IT0u6gQJaOhKnskYsQ3gK-BQR2iaIDGMh0JAPQ4xPqDxd-oLO9eG64Ekw.png 546w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/RfzQSZ0u_6nSruQZiSRaosGwq9Zz5bEnwkoCNdbGTT-IT0u6gQJaOhKnskYsQ3gK-BQR2iaIDGMh0JAPQ4xPqDxd-oLO9eG64Ekw-300x54.png 300w\" sizes=\"(max-width: 546px) 100vw, 546px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Confusion matrix results:<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"411\" height=\"72\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/cTg2t5XjRWq_PKWzMxnwKy2SOBOSq_XIGPSDDnXYD-kr5Q22wmYri558JJlEPQG_TI-zyPbWrqdNvYE6dWEFg3BtSuGBEz66IkWM.png\" alt=\"\" class=\"wp-image-12280\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/cTg2t5XjRWq_PKWzMxnwKy2SOBOSq_XIGPSDDnXYD-kr5Q22wmYri558JJlEPQG_TI-zyPbWrqdNvYE6dWEFg3BtSuGBEz66IkWM.png 411w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/cTg2t5XjRWq_PKWzMxnwKy2SOBOSq_XIGPSDDnXYD-kr5Q22wmYri558JJlEPQG_TI-zyPbWrqdNvYE6dWEFg3BtSuGBEz66IkWM-300x53.png 300w\" sizes=\"(max-width: 411px) 100vw, 411px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Based on the confusion matrix, we proceed to calculate the metrics that indicate the model's quality. We have the following:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">TP = 273,905<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">FP = 2<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">FN = 2<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">TN = 26091<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Accuracy = 0.9999<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The accuracy is very good; however, as noted in the statistical analysis, fraud cases account for 8.7% of the data, resulting in an imbalanced dataset\u2014that is, there is a large amount of data for one class (non-fraud) and very little for the other (fraud). Given this, we proceed to calculate the F1 score, which is the most appropriate metric for this type of case.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Precision = 0.9999<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Recall = 0.9999<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">F1 score = 0.9999<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It can be seen that the F1 score is very close to 1, so we can conclude that the model is of very high quality and will be highly effective in predicting future fraudulent transactions.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"3_Random_Forest\"><\/span>3. Random Forest<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"704\" height=\"290\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/e7G9mleIe8YNzIPFiqOaoIOhJgTslJ90I2pbOKtCA_x0Lal0x0_5JNmN2wGIOXwVIVtTG-t3aiRer0I7DoJt-rAp5TVKlcMDrken.png\" alt=\"\" class=\"wp-image-12279\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/e7G9mleIe8YNzIPFiqOaoIOhJgTslJ90I2pbOKtCA_x0Lal0x0_5JNmN2wGIOXwVIVtTG-t3aiRer0I7DoJt-rAp5TVKlcMDrken.png 704w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/e7G9mleIe8YNzIPFiqOaoIOhJgTslJ90I2pbOKtCA_x0Lal0x0_5JNmN2wGIOXwVIVtTG-t3aiRer0I7DoJt-rAp5TVKlcMDrken-300x124.png 300w\" sizes=\"(max-width: 704px) 100vw, 704px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The results of the confusion matrix are displayed.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"352\" height=\"62\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/2YpaFocVKsWj-P0GDMEDQ0hDrx20qRc2JEcmo5-ZcAmS0vds6zxeDhAUlRiIU3UzHoz4tkjys1bfpNxR3u7PsOepwRtLzvYfob31.png\" alt=\"\" class=\"wp-image-12278\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/2YpaFocVKsWj-P0GDMEDQ0hDrx20qRc2JEcmo5-ZcAmS0vds6zxeDhAUlRiIU3UzHoz4tkjys1bfpNxR3u7PsOepwRtLzvYfob31.png 352w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/2YpaFocVKsWj-P0GDMEDQ0hDrx20qRc2JEcmo5-ZcAmS0vds6zxeDhAUlRiIU3UzHoz4tkjys1bfpNxR3u7PsOepwRtLzvYfob31-300x53.png 300w\" sizes=\"(max-width: 352px) 100vw, 352px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">As you can see, it yields the same result as the decision tree.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"4_K-nearest_neighbors\"><\/span>4. K-nearest neighbors<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"776\" height=\"228\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/7cvYrWWtMx-Y1L2a5aRc8faVDTQuW-qy3uqF-h6sV_1i_DH6CnLn93heTQWBbOdFCA0dtLf1494bAavOPV7YlCFnzhgF8M5LYzpb.png\" alt=\"\" class=\"wp-image-12277\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/7cvYrWWtMx-Y1L2a5aRc8faVDTQuW-qy3uqF-h6sV_1i_DH6CnLn93heTQWBbOdFCA0dtLf1494bAavOPV7YlCFnzhgF8M5LYzpb.png 776w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/7cvYrWWtMx-Y1L2a5aRc8faVDTQuW-qy3uqF-h6sV_1i_DH6CnLn93heTQWBbOdFCA0dtLf1494bAavOPV7YlCFnzhgF8M5LYzpb-300x88.png 300w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/7cvYrWWtMx-Y1L2a5aRc8faVDTQuW-qy3uqF-h6sV_1i_DH6CnLn93heTQWBbOdFCA0dtLf1494bAavOPV7YlCFnzhgF8M5LYzpb-768x226.png 768w\" sizes=\"(max-width: 776px) 100vw, 776px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Confusion matrix results.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"222\" height=\"65\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/Zqd6JJ6yBernplUTPzChEHxhy7drKlqvernt0HiYIbJ0Qk2Ezn0e4trHjFcivC7dvfNN9z8H8NCfS-fVUFvi4rHV81trisxMWJrc.png\" alt=\"\" class=\"wp-image-12276\"\/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The table above shows that the quality of this model is lower than that of the model using the decision tree method. The model's evaluation metrics are calculated as follows:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">TP = 273,713<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">FP = 194<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">FN = 500<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">TN = 25593<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Accuracy = 0.9977<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Precision = 0.9993<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Recall = 0.9982<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">F1 score = 0.9987<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"5_Naive_bayes\"><\/span>5. Naive Bayes<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"467\" height=\"210\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/5gqPXS_gOjYYFUWkJUFR3aggmJwWo0ladB9DUnQHfJbzIc3vjuOHqce61Um1euKkUyhOvKcW5wHjOIg-HJQT08r0_uFCndkscvLE.png\" alt=\"\" class=\"wp-image-12275\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/5gqPXS_gOjYYFUWkJUFR3aggmJwWo0ladB9DUnQHfJbzIc3vjuOHqce61Um1euKkUyhOvKcW5wHjOIg-HJQT08r0_uFCndkscvLE.png 467w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/5gqPXS_gOjYYFUWkJUFR3aggmJwWo0ladB9DUnQHfJbzIc3vjuOHqce61Um1euKkUyhOvKcW5wHjOIg-HJQT08r0_uFCndkscvLE-300x135.png 300w\" sizes=\"(max-width: 467px) 100vw, 467px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Confusion matrix results:<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"296\" height=\"73\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/d3O4nQD7epAiDz8yMhv3IHHGrtledaRJOueyoABfpqGLNTm6pjkaUIk4_pZnowCBMFSwmN80k-hoL5fhF2tt7eF3sf9g_GzI7heL.png\" alt=\"\" class=\"wp-image-12274\"\/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">It can be seen that with this method, the model's quality declines even further compared to the previous models. The evaluation metrics are shown below:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">TP = 269,793<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">FP = 4114<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">FN = 10801<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">TN = 15292<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Accuracy = 0.9503<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Precision = 0.9850<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Recall = 0.9615<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">F1 score = 0.9731<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"_6_Comparacion_de_los_modelos\"><\/span>&nbsp;6. Comparison of the Models<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Finally, as a summary, a comparative table of the metrics for the different models used is provided, showing that the method with the best metrics\u2014and therefore the one that should be used\u2014is the decision tree, with an F1 score of 99.99%.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"140\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/4fSrPCE-_yLKukFl53PIhsJGFeW4i68VC8A5rxxFFzs9Otx0-rCykipOHyvfQ4GtXFfAZRGaFRtchuh70dm1v3Fpwdwqur3wygPJ-1024x140.png\" alt=\"\" class=\"wp-image-12273\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/4fSrPCE-_yLKukFl53PIhsJGFeW4i68VC8A5rxxFFzs9Otx0-rCykipOHyvfQ4GtXFfAZRGaFRtchuh70dm1v3Fpwdwqur3wygPJ-1024x140.png 1024w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/4fSrPCE-_yLKukFl53PIhsJGFeW4i68VC8A5rxxFFzs9Otx0-rCykipOHyvfQ4GtXFfAZRGaFRtchuh70dm1v3Fpwdwqur3wygPJ-300x41.png 300w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/4fSrPCE-_yLKukFl53PIhsJGFeW4i68VC8A5rxxFFzs9Otx0-rCykipOHyvfQ4GtXFfAZRGaFRtchuh70dm1v3Fpwdwqur3wygPJ-768x105.png 768w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/4fSrPCE-_yLKukFl53PIhsJGFeW4i68VC8A5rxxFFzs9Otx0-rCykipOHyvfQ4GtXFfAZRGaFRtchuh70dm1v3Fpwdwqur3wygPJ-1536x209.png 1536w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/4fSrPCE-_yLKukFl53PIhsJGFeW4i68VC8A5rxxFFzs9Otx0-rCykipOHyvfQ4GtXFfAZRGaFRtchuh70dm1v3Fpwdwqur3wygPJ.png 1600w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"7_Clustering\"><\/span><strong>7. Clustering<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Another important aspect of the business is categorizing transactions, by<strong> taking advantage of the benefits offered by machine learning and data mining<\/strong> We then proceed to perform the clustering. For this purpose, we worked with the numerical variables available in the dataset. As can be seen in the visualization (Power BI dashboard), the variables<em> \u201cdistance from home\u201d<\/em> y <em>\u201cratio relative to the average purchase price\u201d<\/em> These were the variables that showed the greatest variation in medians when fraud cases were analyzed separately from non-fraud cases; for this reason, these two variables are used in the clustering process.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The Elbow method is used to determine the optimal number of clusters.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The results of this method are shown below; as you can see, the number of clusters to use is 3.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"818\" height=\"127\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/OreR_qBwQ5QFr2Tos1iB2KqxDgKjysF8VssVbwComqpX0EaXxvv8jnBZSn5Jq_3A7dKMg4vt8CkMQvmfyRPmbpfacw3OkwwRIC7r.png\" alt=\"\" class=\"wp-image-12272\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/OreR_qBwQ5QFr2Tos1iB2KqxDgKjysF8VssVbwComqpX0EaXxvv8jnBZSn5Jq_3A7dKMg4vt8CkMQvmfyRPmbpfacw3OkwwRIC7r.png 818w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/OreR_qBwQ5QFr2Tos1iB2KqxDgKjysF8VssVbwComqpX0EaXxvv8jnBZSn5Jq_3A7dKMg4vt8CkMQvmfyRPmbpfacw3OkwwRIC7r-300x47.png 300w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/OreR_qBwQ5QFr2Tos1iB2KqxDgKjysF8VssVbwComqpX0EaXxvv8jnBZSn5Jq_3A7dKMg4vt8CkMQvmfyRPmbpfacw3OkwwRIC7r-768x119.png 768w\" sizes=\"(max-width: 818px) 100vw, 818px\" \/><\/figure>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"386\" height=\"278\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/IwPsYgJSK-lC0W4VSfS8_M1yIomY0uEkrD6Y8etiSqh2rpOV8fP_JjblRy2DKV2demhU9fVnAvNdB7QQ4HIN7JLfEL-OytBymSC-.png\" alt=\"\" class=\"wp-image-12271\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/IwPsYgJSK-lC0W4VSfS8_M1yIomY0uEkrD6Y8etiSqh2rpOV8fP_JjblRy2DKV2demhU9fVnAvNdB7QQ4HIN7JLfEL-OytBymSC-.png 386w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/IwPsYgJSK-lC0W4VSfS8_M1yIomY0uEkrD6Y8etiSqh2rpOV8fP_JjblRy2DKV2demhU9fVnAvNdB7QQ4HIN7JLfEL-OytBymSC--300x216.png 300w\" sizes=\"(max-width: 386px) 100vw, 386px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Figure 8. WCSS as a function of K (number of clusters).<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Once the optimal number of clusters has been determined, we proceed to perform the clustering and create a graph to visualize the distribution of the data within each cluster:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><img loading=\"lazy\" decoding=\"async\" width=\"150\" height=\"99\" class=\"wp-image-12270\" style=\"width: 150px;\" src=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/ceaYyCFAqLCYj60jUNegNAB9Hu6uLdpNWr1QGSvh4zLjiKE3yJE5DmbapCRAhDtbpg4Zp38c-mAldyWD_ulZg9KMS1Ci3gGd1hfU.png\" alt=\"\" srcset=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/ceaYyCFAqLCYj60jUNegNAB9Hu6uLdpNWr1QGSvh4zLjiKE3yJE5DmbapCRAhDtbpg4Zp38c-mAldyWD_ulZg9KMS1Ci3gGd1hfU.png 375w, https:\/\/blog.structuralia.com\/wp-content\/uploads\/2025\/04\/ceaYyCFAqLCYj60jUNegNAB9Hu6uLdpNWr1QGSvh4zLjiKE3yJE5DmbapCRAhDtbpg4Zp38c-mAldyWD_ulZg9KMS1Ci3gGd1hfU-300x198.png 300w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The graph shows a first cluster (blue) with transactions that are close to home but have a high ratio relative to the average purchase amount. In addition, there is a second cluster (red) with transactions that are at an average distance from home and have an average ratio relative to the average purchase amount. Finally, a third cluster (green) consists of transactions with a low ratio relative to the average purchase price and a high distance from home.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"8_Conclusiones\"><\/span><strong>8. Conclusions<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>The <strong>dataset card_transdata<\/strong> The data is of good quality in its original source, so it requires minimal processing for subsequent steps.<\/li>\n\n\n\n<li>8.7% of the transactions under review are fraudulent.<\/li>\n\n\n\n<li>Result of the <strong>statistical study,<\/strong> It is determined that the variables \u201cdistance from home\u201d and \u201csame retailer\u201d are correlated.<\/li>\n\n\n\n<li>With regard to fraud, it is <strong>determined that all the variables in the dataset<\/strong> are influential, except for the \u201csame retailer\u201d variable.<\/li>\n\n\n\n<li>In the<strong> correlation study<\/strong> It can be seen that the security measures adopted\u2014such as the use of chips and PIN numbers\u2014are effective in preventing fraud, as they show a negative correlation with fraud cases.<\/li>\n\n\n\n<li>Various models are developed using supervised learning methods to detect cases of fraud; the most effective model is the one that uses the decision tree method, with an F1 score of 99.99%; therefore, it is the <strong>the model that should be used.<\/strong><\/li>\n\n\n\n<li>The transactions were categorized into <strong>3 clusters<\/strong>, using the continuous variables \u201cdistance from home\u201c and \u201dratio relative to the average purchase amount,\u201d applying the elbow method to determine the optimal number of clusters, and using the K-Means algorithm for the clustering itself.<\/li>\n<\/ul>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<p class=\"wp-block-paragraph\">The Structuralia team would like to thank Magaly Fonseca for her excellent work. We wish her every success in her professional career and in all the challenges she takes on in the future!<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"RESENA_DEL_AUTOR\"><\/span><strong>AUTHOR'S REVIEW:<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong><\/strong><a href=\"https:\/\/www.linkedin.com\/in\/magaly-fonseca-94309152\/\">Magaly Fonseca Maroto<\/a>, with a bachelor's degree in <strong>Industrial Production Engineering at the Technological Institute of Costa Rica<\/strong>. He also completed a technical certificate in computer networking and has taken courses in Six Sigma Green Belt, Customer Service, and PowerBI. He recently earned a <strong>Master's Degree in Big Data and Business Analytics<\/strong> in <strong>Structuralia.<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">He has been working for 12 years at ICE (Instituto Costarricense de Electricidad), Costa Rica\u2019s state-owned electricity and telecommunications company, where he has led various technical teams and is responsible for preparing management reports for the department and participating in improvement projects.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"TESTIMONIO_DEL_AUTOR\"><\/span><strong>AUTHOR'S STATEMENT:<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>1. Why did you choose Structuralia?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>\u00abI chose Structuralia because it was easy to study online. I found the curriculum to be very comprehensive and well-suited to my needs, and I was also able to apply for an OAS scholarship.\".<\/em><em>\u00ab<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>2. What would you highlight most about the master's program?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>\u00abWhat stands out to me about the Master\u2019s in Big Data and Business Analytics is the breadth of topics covered. As a professional in industrial engineering without a strong background in computer science and programming, the way the program covers these topics allowed me to gain a deeper understanding and apply what I learned. I really appreciate that the methodology focuses on helping students gain a thorough understanding of each topic through a solid theoretical foundation combined with practical application, which is essential for internalizing what is covered in theory.\u201d.<br>\u00bbIn addition, the fact that students can work through the topics individually at their own pace is a major advantage of Structuralia. Combined with the user-friendly platform and the way it tracks the percentage of progress for each module and the master\u2019s program as a whole, it\u2019s very useful.\".<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>3. How has it helped you, or how do you think it could help you, in your current or future professional development?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>\u00abThroughout my professional career, I have come to understand the importance of having access to data quickly, in a timely and effective manner, so that it can be transformed into information that enables robust and sound decision-making. This master\u2019s program has helped me complement my career as an industrial engineer and enhance my skills with knowledge in Big Data and Business Analytics. Since these are rapidly growing fields, I am confident that the knowledge I have gained will open many doors for me to continue growing professionally and contribute to the development of the company where I work.\u00bb.<\/em><\/p>","protected":false},"excerpt":{"rendered":"<p>Find out how a student has developed an innovative model to detect transaction fraud and combat cyberattacks in this master's thesis.<\/p>","protected":false},"author":9,"featured_media":8475,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_ayudawp_aiss_exclude":false,"_ayudawp_aiss_summary":"","_ayudawp_aiss_summary_provider":"","_ayudawp_aiss_summary_hash":"","footnotes":""},"categories":[266],"tags":[],"class_list":["post-8014","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-transformacion-digital"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.8 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Proyecto TFM: Modelo para la detecci\u00f3n de fraudes en transacciones - Blog de Structuralia<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/blog.structuralia.com\/en\/wp-json\/wp\/v2\/posts\/8014\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Proyecto TFM: Modelo para la detecci\u00f3n de fraudes en transacciones - Blog de Structuralia\" \/>\n<meta property=\"og:description\" content=\"Descubre c\u00f3mo un alumno ha desarrollado un modelo innovador para detectar fraudes en transacciones y combatir los ciberataques en este TFM.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones\" \/>\n<meta property=\"og:site_name\" content=\"Blog de Structuralia\" \/>\n<meta property=\"article:published_time\" content=\"2025-04-22T07:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-03-27T12:46:00+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2024\/12\/Proyecto-TF-Magaly-Fonseca.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1050\" \/>\n\t<meta property=\"og:image:height\" content=\"550\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"Equipo de redactores\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Equipo de redactores\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"16 minutes\" \/>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Master's Thesis Project: A Model for Detecting Fraud in Transactions - Structuralia Blog","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/blog.structuralia.com\/en\/wp-json\/wp\/v2\/posts\/8014","og_locale":"en_US","og_type":"article","og_title":"Proyecto TFM: Modelo para la detecci\u00f3n de fraudes en transacciones - Blog de Structuralia","og_description":"Descubre c\u00f3mo un alumno ha desarrollado un modelo innovador para detectar fraudes en transacciones y combatir los ciberataques en este TFM.","og_url":"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones","og_site_name":"Blog de Structuralia","article_published_time":"2025-04-22T07:00:00+00:00","article_modified_time":"2026-03-27T12:46:00+00:00","og_image":[{"width":1050,"height":550,"url":"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2024\/12\/Proyecto-TF-Magaly-Fonseca.jpg","type":"image\/jpeg"}],"author":"Equipo de redactores","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Equipo de redactores","Est. reading time":"16 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones#article","isPartOf":{"@id":"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones"},"author":{"name":"Equipo de redactores","@id":"https:\/\/blog.structuralia.com\/#\/schema\/person\/87ddd066f2b0e01e832284f797138b7b"},"headline":"Proyecto TFM: Modelo para la detecci\u00f3n de fraudes en transacciones","datePublished":"2025-04-22T07:00:00+00:00","dateModified":"2026-03-27T12:46:00+00:00","mainEntityOfPage":{"@id":"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones"},"wordCount":3143,"commentCount":0,"publisher":{"@id":"https:\/\/blog.structuralia.com\/#organization"},"image":{"@id":"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones#primaryimage"},"thumbnailUrl":"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2024\/12\/Proyecto-TF-Magaly-Fonseca.jpg","articleSection":["Transformaci\u00f3n digital"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones#respond"]}]},{"@type":"WebPage","@id":"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones","url":"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones","name":"Master's Thesis Project: A Model for Detecting Fraud in Transactions - Structuralia Blog","isPartOf":{"@id":"https:\/\/blog.structuralia.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones#primaryimage"},"image":{"@id":"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones#primaryimage"},"thumbnailUrl":"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2024\/12\/Proyecto-TF-Magaly-Fonseca.jpg","datePublished":"2025-04-22T07:00:00+00:00","dateModified":"2026-03-27T12:46:00+00:00","breadcrumb":{"@id":"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones#primaryimage","url":"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2024\/12\/Proyecto-TF-Magaly-Fonseca.jpg","contentUrl":"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2024\/12\/Proyecto-TF-Magaly-Fonseca.jpg","width":1050,"height":550,"caption":"Proyecto TFM Magaly Fonseca"},{"@type":"BreadcrumbList","@id":"https:\/\/blog.structuralia.com\/en\/proyecto-tfm-modelo-para-la-deteccion-de-fraudes-en-transacciones#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Inicio","item":"https:\/\/blog.structuralia.com"},{"@type":"ListItem","position":2,"name":"Transformaci\u00f3n digital","item":"https:\/\/blog.structuralia.com\/transformacion-digital"},{"@type":"ListItem","position":3,"name":"Proyecto TFM: Modelo para la detecci\u00f3n de fraudes en transacciones"}]},{"@type":"WebSite","@id":"https:\/\/blog.structuralia.com\/#website","url":"https:\/\/blog.structuralia.com\/","name":"Structuralia's Official Blog","description":"On our blog, we create specialized content to keep you up to date on all the latest news, trends, and tips related to engineering.","publisher":{"@id":"https:\/\/blog.structuralia.com\/#organization"},"alternateName":"STRUCTURALIA","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/blog.structuralia.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/blog.structuralia.com\/#organization","name":"Structuralia Blog","url":"https:\/\/blog.structuralia.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/blog.structuralia.com\/#\/schema\/logo\/image\/","url":"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2026\/01\/logo.svg","contentUrl":"https:\/\/blog.structuralia.com\/wp-content\/uploads\/2026\/01\/logo.svg","width":36,"height":38,"caption":"Blog de Structuralia"},"image":{"@id":"https:\/\/blog.structuralia.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/blog.structuralia.com\/#\/schema\/person\/87ddd066f2b0e01e832284f797138b7b","name":"Editorial Team","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/6047234e84771ae0815c2fc62265d85296754868378b6dd88865366408801df4?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/6047234e84771ae0815c2fc62265d85296754868378b6dd88865366408801df4?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/6047234e84771ae0815c2fc62265d85296754868378b6dd88865366408801df4?s=96&d=mm&r=g","caption":"Equipo de redactores"},"url":"https:\/\/blog.structuralia.com\/en\/autor\/equipo-de-redactores"}]}},"views":471,"_links":{"self":[{"href":"https:\/\/blog.structuralia.com\/en\/wp-json\/wp\/v2\/posts\/8014","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.structuralia.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.structuralia.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.structuralia.com\/en\/wp-json\/wp\/v2\/users\/9"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.structuralia.com\/en\/wp-json\/wp\/v2\/comments?post=8014"}],"version-history":[{"count":0,"href":"https:\/\/blog.structuralia.com\/en\/wp-json\/wp\/v2\/posts\/8014\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/blog.structuralia.com\/en\/wp-json\/wp\/v2\/media\/8475"}],"wp:attachment":[{"href":"https:\/\/blog.structuralia.com\/en\/wp-json\/wp\/v2\/media?parent=8014"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.structuralia.com\/en\/wp-json\/wp\/v2\/categories?post=8014"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.structuralia.com\/en\/wp-json\/wp\/v2\/tags?post=8014"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}