<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE ArticleSet PUBLIC "-//NLM//DTD PubMed 2.7//EN" "https://dtd.nlm.nih.gov/ncbi/pubmed/in/PubMed.dtd">
<ArticleSet>
<Article>
<Journal>
				<PublisherName>دانشگاه صنعتی شریف</PublisherName>
				<JournalTitle>مهندسی صنایع و مدیریت</JournalTitle>
				<Issn>2676-4741</Issn>
				<Volume>41</Volume>
				<Issue>2</Issue>
				<PubDate PubStatus="epublish">
					<Year>2026</Year>
					<Month>03</Month>
					<Day>20</Day>
				</PubDate>
			</Journal>
<ArticleTitle>Topic Modeling and Sentiment Analysis of Customers Using Natural Language Processing and Machine Learning Techniques</ArticleTitle>
<VernacularTitle>مدل‌سازی موضوعی و تحلیل احساس‌های مشتریان با استفاده از روش‌های پردازش زبان طبیعی و یادگیری ماشین</VernacularTitle>
			<FirstPage>45</FirstPage>
			<LastPage>61</LastPage>
			<ELocationID EIdType="pii">24119</ELocationID>
			
<ELocationID EIdType="doi">10.24200/j65.2025.65318.2418</ELocationID>
			
			<Language>FA</Language>
<AuthorList>
<Author>
					<FirstName>راضیه</FirstName>
					<LastName>رحیمی</LastName>
<Affiliation>دانشکده مهندسی نساجی، دانشگاه صنعتی امیرکبیر، تهران، ایران</Affiliation>

</Author>
<Author>
					<FirstName>رضا</FirstName>
					<LastName>قاسمی یقین</LastName>
<Affiliation>دانشکده‌ی مهندسی صنایع و سیستم‌های مدیریت، دانشگاه صنعتی امیرکبیر، تهران، ایران.</Affiliation>

</Author>
</AuthorList>
				<PublicationType>Journal Article</PublicationType>
			<History>
				<PubDate PubStatus="received">
					<Year>2024</Year>
					<Month>10</Month>
					<Day>15</Day>
				</PubDate>
			</History>
		<Abstract>Given that modeling and predicting customer behavior using data science helps companies gain a better understanding of customer behavior, this research focuses on analyzing customer reviews in the women’s clothing domain within e-commerce. We employ machine learning techniques and natural language processing (NLP) to achieve this goal. The machine learning models used include Support Vector Machine, Logistic Regression, Decision Tree, Random Forest, Multinomial Naive Bayes, Complement Naive Bayes, XGBoost, and LightGBM. To extract and vectorize text features from the reviews, we utilize the TF-IDF and Word2vec algorithms. We employ Topic Modeling using the Latent Dirichlet Allocation (LDA) method and k-means clustering. The dataset consists of women’s clothing reviews, with the target variable being customer ratings in those reviews. The study is conducted in binary, three-class, and five-class scenarios. The target variable, which originally has five classes (scores 1 to 5), is categorized into two-class and three-class modes. In the two-class mode, scores below 3 are class zero, while scores of 3 and above are class one. In the three-class mode, scores below 3 are class zero, scores equal to 3 are class one, and scores above 3 are class two. In all three cases, the Random Forest model performs best, achieving an accuracy of 0.98 in the binary case, 0.95 in the three-class case, and 0.91 in the five-class case. After performing the required preprocessing and feature engineering, principal component analysis (PCA) and T SNE are applied. After that, the scatter diagram of the data is drawn and the optimal number of clusters, 3, is estimated using the elbow and silhouette diagrams. In the next step, by removing punctuation marks, stop words and words with fewer than three letters, converting the first letter of the words to lowercase and lemmatization, data cleaning was done. After that, topic modeling was done and each of the topics and words related to them were examined. In the next step, the topics were examined in different clusters.   These analyses provide a comprehensive understanding of the key themes and concerns customers have when considering womenswear items in each of the four topics.</Abstract>
			<OtherAbstract Language="FA">در نوشتارحاضر، به بررسی و تجزیه­و­تحلیل نظرهای مشتریان پوشاک زنان با استفاده از روش­های یادگیری ماشین و پردازش زبان طبیعی پرداخته­ شده است. مدل­های یادگیری ماشین استفاده‌شده، شامل: ماشین بردار پشتیبان، رگرسیون لجستیک، درخت تصمیم، جنگل تصادفی، بیز ساده‌ی چندجمله­ای، بیز ساده‌ی مکمل، XGBoost ، و LightGBM بوده‌اند. به‌منظور برداری‌کردن و استخراج ویژگی متن نظرها از الگوریتم‌های TF-IDF و Word2vec استفاده شده است. مدل­سازی موضوعی با استفاده از روش تخصیص دیریکله‌ی پنهان و خوشه­بندی K-Means انجام شده است. مجموعه‌ی داده‌ی استفاده‌شده مربوط به نظرهای پوشاک زنان و متغیر هدف، امتیازهای مشتریان در نظرها بوده است. مطالعه‌ی حاضر، در حالت‌های: دو­کلاسه، سه­کلاسه، و پنج­کلاسه انجام و در هر سه حالت، بهترین عملکرد مربوط به مدل جنگل تصادفی با دقت 98&lt;sub&gt;/&lt;/sub&gt;0 در حالت دوکلاسه، 95&lt;sub&gt;/&lt;/sub&gt;0 در حالت سه­کلاسه، و 91&lt;sub&gt;/&lt;/sub&gt;0 در حالت پنج­کلاسه برآورد شده است.</OtherAbstract>
		<ObjectList>
			<Object Type="keyword">
			<Param Name="value">تجزیه‌و‌تحلیل احساس‌ها</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">یادگیری ماشین</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">پردازش زبان طبیعی</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">متن‌کاوی</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">رفتار مشتری</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">مدل‌سازی موضوعی</Param>
			</Object>
		</ObjectList>
<ArchiveCopySource DocType="pdf">https://sjie.journals.sharif.edu/article_24119_883591bae75b035f5a2e0fbab701660b.pdf</ArchiveCopySource>
</Article>
</ArticleSet>
