<?xml version="1.0" encoding="UTF-8"?>
    <!DOCTYPE article PUBLIC "-//NLM/DTD JATS (Z39.96) Journal Publishing DTD v1.2 20120330//EN" "http://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
    <!--<?xml-stylesheet type="text/xsl" href="article.xsl">-->
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:ns1="http://www.w3.org/1999/xlink" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.2" xml:lang="en">
	<front>
		<journal-meta>
			<journal-id journal-id-type="issn">2303-9868</journal-id>
			<journal-id journal-id-type="eissn">2227-6017</journal-id>
			<journal-title-group>
				<journal-title>International Research Journal</journal-title>
			</journal-title-group>
			<issn pub-type="epub">2303-9868</issn>
			<publisher>
				<publisher-name>Cifra LLC</publisher-name>
			</publisher>
		</journal-meta>
		<article-meta>
			<article-id pub-id-type="doi">10.60797/IRJ.2026.170.51</article-id>
			<article-categories>
				<subj-group>
					<subject>Brief communication</subject>
				</subj-group>
			</article-categories>
			<title-group>
				<article-title>Computer-Vision Recognition of Unsafe Actions and Personal Protective Equipment Violations from a Video Stream</article-title>
			</title-group>
			<contrib-group>
				<contrib contrib-type="author" corresp="yes">
					<name>
						<surname>Gayfutdinov</surname>
						<given-names>Timur Ildarovich</given-names>
					</name>
					<email>odie@list.ru</email>
					<xref ref-type="aff" rid="aff-1">1</xref>
				</contrib>
				<contrib contrib-type="author">
					<name>
						<surname>Medvedev</surname>
						<given-names>Pavel Sergeevich</given-names>
					</name>
					<email>pasha97_1997@mail.ru</email>
					<xref ref-type="aff" rid="aff-1">1</xref>
				</contrib>
				<contrib contrib-type="author">
					<contrib-id contrib-id-type="rinc">https://elibrary.ru/author_profile.asp?id=706627</contrib-id>
					<name>
						<surname>Mokshin</surname>
						<given-names>Vladimir Vasilevich</given-names>
					</name>
					<email>vmk550te@mail.ru</email>
					<xref ref-type="aff" rid="aff-1">1</xref>
				</contrib>
			</contrib-group>
			<aff id="aff-1">
				<label>1</label>
				<institution>Kazan National Research Technical University named after A.N. Tupolev – KAI</institution>
			</aff>
			<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2026-08-17">
				<day>17</day>
				<month>08</month>
				<year>2026</year>
			</pub-date>
			<pub-date pub-type="collection">
				<year>2026</year>
			</pub-date>
			<volume>8</volume>
			<issue>170</issue>
			<fpage>1</fpage>
			<lpage>8</lpage>
			<history>
				<date date-type="received" iso-8601-date="2026-06-06">
					<day>06</day>
					<month>06</month>
					<year>2026</year>
				</date>
				<date date-type="accepted" iso-8601-date="2026-08-04">
					<day>04</day>
					<month>08</month>
					<year>2026</year>
				</date>
			</history>
			<permissions>
				<copyright-statement>Copyright: &amp;#x00A9; 2022 The Author(s)</copyright-statement>
				<copyright-year>2022</copyright-year>
				<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
					<license-p>
						This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. See 
						<uri xlink:href="http://creativecommons.org/licenses/by/4.0/">http://creativecommons.org/licenses/by/4.0/</uri>
					</license-p>
					.
				</license>
			</permissions>
			<self-uri xlink:href="https://research-journal.org/archive/8-170-2026-august/10.60797/IRJ.2026.170.51"/>
			<abstract>
				<p>Workplace injury is largely caused by failure to comply with personal protective equipment (PPE) requirements and by violations of work procedures. This paper investigates the use of computer vision methods for the automatic detection of such violations from the video stream of industrial cameras. An information system is proposed and implemented whose core is a fine-tuned YOLOv5 convolutional neural network detector. The system not only records the presence or absence of a safety helmet and a high-visibility vest on a worker but also recognizes a dangerous action — climbing into the bed of a truck without a step ladder — which requires accounting for the spatial context of the scene. The decision logic is formalized as an interpretable safety-violation factor that combines per-object detection confidences, the spatial relations among the worker, the protective items and the vehicle, and the temporal evidence accumulated along object tracks. A Kalman filter is employed for robust object tracking across frames and for smoothing trajectories, and detected violations are automatically compiled into reports. The composition of the training set, the classes of recognized objects, the data preparation and versioning procedure, the training parameters, and the quality-assessment metrics are described. An example of the system operating in an active production workshop is presented, and the limitations of the approach — related primarily to the volume and resolution of the training data — are discussed.</p>
			</abstract>
			<kwd-group>
				<kwd>computer vision</kwd>
				<kwd> object detection</kwd>
				<kwd> YOLOv5</kwd>
				<kwd> convolutional neural networks</kwd>
				<kwd> occupational safety</kwd>
				<kwd> personal protective equipment</kwd>
				<kwd> video analytics</kwd>
				<kwd> Kalman filter</kwd>
				<kwd> violation monitoring</kwd>
				<kwd> safety-violation factor</kwd>
				<kwd> industrial safety</kwd>
			</kwd-group>
		</article-meta>
	</front>
	<body>
		<sec>
			<title>HTML-content</title>
			<p>1. Introduction</p>
			<p>Workplace injury remains a pressing problem, and a significant proportion of accidents at enterprises is associated with workers neglecting personal protective equipment (PPE). This proportion can be reduced by a system that would monitor PPE use in real-time and record violations. The present work is devoted to the development of such a system.</p>
			<p>A substantial share of industrial accidents is preventable and is caused not by equipment failure but by the human factor — a refusal to wear a safety helmet or a high-visibility vest, a breach of the procedure for loading operations, or the neglect of step ladders and guard rails. Traditional monitoring of occupational-safety compliance relies on periodic walk-throughs and the selective review of surveillance recordings; that is, it is episodic in nature and does not allow a response to a violation at the moment it occurs. Meanwhile, the ubiquity of video-surveillance systems and the maturity of deep-learning algorithms create the preconditions for continuous automatic monitoring capable of detecting violations in real time and thereby reducing the risk of injury.</p>
			<p>Almost every production process involves the use of PPE, and enterprises build entire sets of measures to ensure occupational-safety compliance. Nevertheless, some workers neglect these requirements and fail to put on a high-visibility vest or a helmet. A typical situation is telling: in order to load a part more quickly, a person climbs into the truck bed straight from the ground instead of positioning a ladder as the rules prescribe. The price of such liberties is injuries and accidents, so detecting them in a timely manner is fundamentally important.</p>
			<p>The system is written in Python and uses a convolutional neural network detector of the YOLOv5 family, fine-tuned to the conditions of a specific enterprise by means of transfer learning: weights pretrained on a large publicly available image set are taken as the starting point, after which the network is adapted to the target classes on the enterprise’s own data. The choice of the YOLO family is dictated by its combination of high-speed and acceptable accuracy: these models perform well in recognition tasks in manufacturing, transport, and other domains where a decision must be made without delay.</p>
			<p>The aim of this work is to investigate computer vision methods as applied to the task of automatically detecting occupational-safety violations and, on this basis, to develop a workable video-analytics system. To achieve this aim, the following tasks were addressed: a comparative analysis of modern object-detection architectures and a justification for the choice of base model; the creation and annotation of a dataset reflecting characteristic situations of a production workshop; fine-tuning of the detector model on the target classes; implementation of a mechanism for tracking objects across frames; organization of the automatic logging of detected violations; and assessment of the quality of the resulting solution.</p>
			<p>Most existing systems for monitoring personal protective equipment treat the task as the per-frame detection of isolated objects and reduce a violation to the mere absence of a bounding box, without a formal rule that turns the detector outputs into a decision and without explicit treatment of the temporal instability of those outputs. The scientific novelty of this work is accordingly twofold. First, a formalized model of the safety-violation factor is proposed: the factor is an interpretable scalar variable that maps the raw detector outputs — class labels, bounding boxes and confidence scores — into a graded measure of danger by combining a confidence criterion, the pairwise spatial relations among the worker, the protective items and the vehicle, and a recursive temporal aggregation of the evidence along an object track. The model supplies a reproducible decision rule with explicit, tunable parameters in place of ad hoc heuristics, and it is what enables the system to recognize not only the presence or absence of individual protective items but also a dangerous action defined by the spatial context of the scene — a worker climbing into the bed of a truck without a step ladder. Second, a complete and reproducible technological pipeline is built around this model — from versioned data preparation and detector fine-tuning to track-level violation scoring and the automatic generation of reports — adapted to the conditions of a specific enterprise.</p>
			<p>2. Review
of existing approaches</p>
			<p>Modern object detectors are commonly divided into two-stage detectors, in which candidate regions are first generated and then classified, and one-stage detectors, which perform localization and classification in a single pass of the network. For real-time systems, processing latency is critical, so it is precisely the comparison of these two approaches that determines the choice of architecture.</p>
			<p>Among object-detection approaches, the R-CNN family with its modifications and the YOLO algorithm stand out. The original R-CNN first selects candidate regions by selective search and then classifies each of them. This two-stage scheme rests on a fixed selection algorithm and proves too slow for real-time operation.</p>
			<p>Fast R-CNN removes the principal source of latency: the entire image passes through the convolutional network only once, and candidate regions are projected onto a shared feature map by the RoI Pooling operation, so that the convolution for each region is no longer recomputed </p>
			<p>[1]</p>
			<p>Faster R-CNN goes further and abandons slow selective search altogether: candidate regions are now proposed by a trainable region proposal network (RPN) </p>
			<p>[2]</p>
			<p>YOLO is among the fastest detectors: the entire scene is analyzed in a single pass of the network, without a separate region-generation stage. The comparison in </p>
			<p>[3][4][5][6]</p>
			<p>The YOLO family has continued to develop well beyond the version used here. Recent releases move the accuracy–latency frontier through contrasting design philosophies. YOLO12 (2025) adopts an attention-centric architecture, introducing an area-attention mechanism and residual efficient layer-aggregation networks that raise detection accuracy at a comparable parameter budget, at the cost of heavier memory use on the central processor </p>
			<p>[7][8]</p>
			<p>3. Proposed
method</p>
			<p>The basic YOLOv5 architecture follows the “backbone — neck — head” principle: the backbone extracts image features, the neck aggregates them at several scales, and the head produces the final predictions. Having received an image, the network builds feature maps at several scales and connects the levels to one another via skip connections </p>
			<p>[9][10][11][12]</p>
			<p>YOLOv5 itself is a modern detection model implemented in the PyTorch framework </p>
			<p>[13][14]</p>
			<p>For each potential box, the network outputs five numbers — the center coordinates x and y, the width, the height, and the confidence (objectness) score </p>
			<p>[15][16][17][18]</p>
			<p>Two engineering decisions noticeably affected the training quality of this model. The first is the Focus input layer, which in YOLOv5 replaced the starting block of several convolutions used in earlier versions </p>
			<p>[19][20][21][22]</p>
			<p>The training set was formed from surveillance frames typical of a production workshop: workers with and without protective clothing, freight vehicles, loading and unloading operations, background equipment, and stored parts. The following object classes were annotated: worker, safety helmet (helmet), high-visibility vest (vest), truck (kamaz), and a dangerous action — climbing into the truck bed without a step ladder (no_stairs). The violation-detection logic is built on relating these classes: if a tracked worker is not associated with a helmet or vest detection, a PPE-use violation is recorded; and if the worker’s region coincides with the truck bed in the absence of a ladder, a violation of the work procedure is registered.</p>
			<p>All images were obtained from the enterprise’s own stationary surveillance cameras rather than from public collections, so the set reproduces exactly the viewing geometry in which the system is intended to operate. The recordings were made by three fixed IP cameras with a native frame resolution of 1920×1080 pixels, mounted at a height of four to six meters and inclined at 25–40 degrees to the horizontal; frames were extracted from the recordings at intervals of at least two seconds and screened for near-duplicates, so that the set does not contain adjacent frames of one and the same scene. The distance from a camera to a worker ranges from about five to twenty-five meters, at which a helmet or a vest occupies from roughly ten to fifty pixels along its larger side in the source frame and becomes smaller still after scaling to the network input — this viewing geometry is the root of the difficulty of the small classes discussed below. The shooting conditions span the day and evening shifts: the workshop is lit by overhead luminaires at all times, daylight from the gates and the window strips is added during the day shift, and the set deliberately includes backlit scenes against open gates, glare on metal surfaces, and shadows cast by the equipment. A typical frame contains from one to four workers, often partially occluded by machinery, stored parts, or the sides of the truck bed; the clothing worn under the vest varies with the season, which adds appearance variability to the worker class.</p>
			<p>The distribution of the images and of the annotated object instances over the classes is summarized in Table 1. The imbalance visible there is a property of the domain rather than of the sampling: a worker is present in almost every frame and is often not alone; protective items are somewhat rarer, because it is precisely their absence that constitutes a violation of interest; the truck appears only in loading scenes; and the dangerous action no_stairs is rare because the violation itself is rare, and staging it deliberately for the camera was ruled out for safety reasons. The remaining frames without people serve as background examples that reduce false positives. The set was divided into training and validation subsets in a fixed 4 : 1 proportion — 800 and 200 images — with stratification by the rarest class, so that the share of no_stairs instances in the validation subset matches its share in the set as a whole; frames extracted from the same recording were always assigned to the same subset, which prevents near-identical scenes from leaking across the split. Table 1 also makes the connection with the per-class quality metrics explicit: the best-recognized class (kamaz) is a large, low-variability object, whereas the weakest (no_stairs) combines the smallest number of examples with the highest variability of pose and viewpoint.</p>
			<table-wrap id="T1">
				<label>Table 1</label>
				<caption>
					<p>Composition of the dataset: images and annotated object instances by class</p>
				</caption>
				<table>
					<tr>
						<td>Class</td>
						<td>Images with the class</td>
						<td>Instances, training</td>
						<td>Instances, validation</td>
						<td>Share of instances, %</td>
					</tr>
					<tr>
						<td>Worker (worker)</td>
						<td>962</td>
						<td>1684</td>
						<td>428</td>
						<td>37.0</td>
					</tr>
					<tr>
						<td>Helmet (helmet)</td>
						<td>874</td>
						<td>1198</td>
						<td>306</td>
						<td>26.3</td>
					</tr>
					<tr>
						<td>Vest (vest)</td>
						<td>843</td>
						<td>1114</td>
						<td>283</td>
						<td>24.5</td>
					</tr>
					<tr>
						<td>Truck (kamaz)</td>
						<td>517</td>
						<td>414</td>
						<td>103</td>
						<td>9.1</td>
					</tr>
					<tr>
						<td>Climbing without a ladder (no_stairs)</td>
						<td>178</td>
						<td>142</td>
						<td>36</td>
						<td>3.1</td>
					</tr>
					<tr>
						<td>Total</td>
						<td>1000</td>
						<td>4552</td>
						<td>1156</td>
						<td>100.0</td>
					</tr>
				</table>
			</table-wrap>
			<p>Data preparation is split off into a separate step that precedes training </p>
			<p>[11][3]</p>
			<p>The model was trained on the set prepared in this way </p>
			<p>[12]</p>
			<p>To track objects robustly from frame to frame and to smooth their trajectories, a Kalman filter was added to the system. It is called linear because its state and measurement models are specified by linear relationships, and each new estimate is constructed recursively — from the previous estimate and the fresh measurement supplied by the detector </p>
			<p>[18]</p>
			<p>We now formalize the violation-detection logic outlined above. At each frame t the detector returns a set of detections; an individual detection is a triple d = (c, b, s) consisting of a class label c, a bounding box b and a confidence score s. We write P with a class superscript for the subsets of detections that are helmets (H), vests (V), the dangerous action no_stairs (A) and trucks (T), and we denote a tracked worker by w. The spatial relation between two boxes is measured by their coverage:</p>
			<mml:math display="inline">
				<mml:mrow>
					<mml:mo>cov</mml:mo>
					<mml:mrow>
						<mml:mo stretchy="true" fence="true" form="prefix">(</mml:mo>
						<mml:msub>
							<mml:mi>b</mml:mi>
							<mml:mrow>
								<mml:mn>1</mml:mn>
							</mml:mrow>
						</mml:msub>
						<mml:mo>,</mml:mo>
						<mml:msub>
							<mml:mi>b</mml:mi>
							<mml:mrow>
								<mml:mn>2</mml:mn>
							</mml:mrow>
						</mml:msub>
						<mml:mo stretchy="true" fence="true" form="postfix">)</mml:mo>
					</mml:mrow>
					<mml:mo>=</mml:mo>
					<mml:mfrac>
						<mml:mrow>
							<mml:mo>area</mml:mo>
							<mml:mrow>
								<mml:mo stretchy="true" fence="true" form="prefix">(</mml:mo>
								<mml:msub>
									<mml:mi>b</mml:mi>
									<mml:mrow>
										<mml:mn>1</mml:mn>
									</mml:mrow>
								</mml:msub>
								<mml:mo>∩</mml:mo>
								<mml:msub>
									<mml:mi>b</mml:mi>
									<mml:mrow>
										<mml:mn>2</mml:mn>
									</mml:mrow>
								</mml:msub>
								<mml:mo stretchy="true" fence="true" form="postfix">)</mml:mo>
							</mml:mrow>
						</mml:mrow>
						<mml:mrow>
							<mml:mo>area</mml:mo>
							<mml:mrow>
								<mml:mo stretchy="true" fence="true" form="prefix">(</mml:mo>
								<mml:msub>
									<mml:mi>b</mml:mi>
									<mml:mrow>
										<mml:mn>1</mml:mn>
									</mml:mrow>
								</mml:msub>
								<mml:mo stretchy="true" fence="true" form="postfix">)</mml:mo>
							</mml:mrow>
						</mml:mrow>
					</mml:mfrac>
				</mml:mrow>
			</mml:math>
			<p>which gives the fraction of the area of the first box that lies inside the second; unlike the symmetric intersection-over-union, this asymmetric quantity is the natural measure of a containment relation, such as a small helmet box lying inside a larger worker box. A protective item of class κ — a helmet or a vest — is regarded as belonging to the worker when its detection is sufficiently confident, sufficiently contained in the worker region, and geometrically plausible:</p>
			<mml:math display="inline">
				<mml:mrow>
					<mml:msub>
						<mml:mi>a</mml:mi>
						<mml:mrow>
							<mml:mi>κ</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo stretchy="false">(</mml:mo>
					<mml:mi>w</mml:mi>
					<mml:mo>,</mml:mo>
					<mml:mi>p</mml:mi>
					<mml:mo stretchy="false">)</mml:mo>
					<mml:mo>=</mml:mo>
					<mml:mrow>
						<mml:mn mathvariant="bold">1</mml:mn>
					</mml:mrow>
					<mml:mrow>
						<mml:mo stretchy="true" fence="true" form="prefix">[</mml:mo>
						<mml:msub>
							<mml:mi>s</mml:mi>
							<mml:mrow>
								<mml:mi>p</mml:mi>
							</mml:mrow>
						</mml:msub>
						<mml:mo>≥</mml:mo>
						<mml:msub>
							<mml:mi>τ</mml:mi>
							<mml:mrow>
								<mml:mi>κ</mml:mi>
							</mml:mrow>
						</mml:msub>
						<mml:mo stretchy="true" fence="true" form="postfix">]</mml:mo>
					</mml:mrow>
					<mml:mi>·</mml:mi>
					<mml:mrow>
						<mml:mn mathvariant="bold">1</mml:mn>
					</mml:mrow>
					<mml:mrow>
						<mml:mo stretchy="true" fence="true" form="prefix">[</mml:mo>
						<mml:mo>cov</mml:mo>
						<mml:mrow>
							<mml:mo stretchy="true" fence="true" form="prefix">(</mml:mo>
							<mml:msub>
								<mml:mi>b</mml:mi>
								<mml:mrow>
									<mml:mi>p</mml:mi>
								</mml:mrow>
							</mml:msub>
							<mml:mo>,</mml:mo>
							<mml:msub>
								<mml:mi>b</mml:mi>
								<mml:mrow>
									<mml:mi>w</mml:mi>
								</mml:mrow>
							</mml:msub>
							<mml:mo stretchy="true" fence="true" form="postfix">)</mml:mo>
						</mml:mrow>
						<mml:mo>≥</mml:mo>
						<mml:msub>
							<mml:mi>θ</mml:mi>
							<mml:mrow>
								<mml:mi>κ</mml:mi>
							</mml:mrow>
						</mml:msub>
						<mml:mo stretchy="true" fence="true" form="postfix">]</mml:mo>
					</mml:mrow>
					<mml:mi>·</mml:mi>
					<mml:msub>
						<mml:mi>g</mml:mi>
						<mml:mrow>
							<mml:mi>κ</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mrow>
						<mml:mo stretchy="true" fence="true" form="prefix">(</mml:mo>
						<mml:msub>
							<mml:mi>b</mml:mi>
							<mml:mrow>
								<mml:mi>p</mml:mi>
							</mml:mrow>
						</mml:msub>
						<mml:mo>,</mml:mo>
						<mml:msub>
							<mml:mi>b</mml:mi>
							<mml:mrow>
								<mml:mi>w</mml:mi>
							</mml:mrow>
						</mml:msub>
						<mml:mo stretchy="true" fence="true" form="postfix">)</mml:mo>
					</mml:mrow>
				</mml:mrow>
			</mml:math>
			<p>Here τ and θ are the per-class confidence and coverage thresholds, and the geometric-consistency term equals one when the centroid of the item lies in the expected sub-region of the worker box — the upper part for a helmet, the torso region for a vest — and zero otherwise. The presence of a helmet and of a vest for the worker at frame t is then the best associated detection of each:</p>
			<mml:math display="inline">
				<mml:mrow>
					<mml:msub>
						<mml:mi>h</mml:mi>
						<mml:mrow>
							<mml:mi>t</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo stretchy="false">(</mml:mo>
					<mml:mi>w</mml:mi>
					<mml:mo stretchy="false">)</mml:mo>
					<mml:mo>=</mml:mo>
					<mml:msub>
						<mml:mo>max</mml:mo>
						<mml:mrow>
							<mml:mi>p</mml:mi>
							<mml:mo>∈</mml:mo>
							<mml:msubsup>
								<mml:mi>P</mml:mi>
								<mml:mrow>
									<mml:mi>t</mml:mi>
								</mml:mrow>
								<mml:mrow>
									<mml:mi>H</mml:mi>
								</mml:mrow>
							</mml:msubsup>
						</mml:mrow>
					</mml:msub>
					<mml:msub>
						<mml:mi>a</mml:mi>
						<mml:mrow>
							<mml:mi>H</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo stretchy="false">(</mml:mo>
					<mml:mi>w</mml:mi>
					<mml:mo>,</mml:mo>
					<mml:mi>p</mml:mi>
					<mml:mo stretchy="false">)</mml:mo>
					<mml:mo>,</mml:mo>
					<mml:mspace width="1em"/>
					<mml:msub>
						<mml:mi>v</mml:mi>
						<mml:mrow>
							<mml:mi>t</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo stretchy="false">(</mml:mo>
					<mml:mi>w</mml:mi>
					<mml:mo stretchy="false">)</mml:mo>
					<mml:mo>=</mml:mo>
					<mml:msub>
						<mml:mo>max</mml:mo>
						<mml:mrow>
							<mml:mi>p</mml:mi>
							<mml:mo>∈</mml:mo>
							<mml:msubsup>
								<mml:mi>P</mml:mi>
								<mml:mrow>
									<mml:mi>t</mml:mi>
								</mml:mrow>
								<mml:mrow>
									<mml:mi>V</mml:mi>
								</mml:mrow>
							</mml:msubsup>
						</mml:mrow>
					</mml:msub>
					<mml:msub>
						<mml:mi>a</mml:mi>
						<mml:mrow>
							<mml:mi>V</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo stretchy="false">(</mml:mo>
					<mml:mi>w</mml:mi>
					<mml:mo>,</mml:mo>
					<mml:mi>p</mml:mi>
					<mml:mo stretchy="false">)</mml:mo>
				</mml:mrow>
			</mml:math>
			<p>with the maximum over an empty set defined as zero, so that a zero value signals a missing item. The dangerous action is contextual: a no_stairs detection is counted only when it is at once attached to the worker and co-located with a vehicle, which expresses the requirement that the worker be climbing into a truck bed rather than merely standing beside it:</p>
			<mml:math display="inline">
				<mml:mrow>
					<mml:msub>
						<mml:mi>n</mml:mi>
						<mml:mrow>
							<mml:mi>t</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo stretchy="false">(</mml:mo>
					<mml:mi>w</mml:mi>
					<mml:mo stretchy="false">)</mml:mo>
					<mml:mo>=</mml:mo>
					<mml:msub>
						<mml:mo>max</mml:mo>
						<mml:mrow>
							<mml:mi>q</mml:mi>
							<mml:mo>∈</mml:mo>
							<mml:msubsup>
								<mml:mi>P</mml:mi>
								<mml:mrow>
									<mml:mi>t</mml:mi>
								</mml:mrow>
								<mml:mrow>
									<mml:mi>A</mml:mi>
								</mml:mrow>
							</mml:msubsup>
						</mml:mrow>
					</mml:msub>
					<mml:mrow>
						<mml:mo stretchy="true" fence="true" form="prefix">(</mml:mo>
						<mml:mrow>
							<mml:mn mathvariant="bold">1</mml:mn>
						</mml:mrow>
						<mml:mrow>
							<mml:mo stretchy="true" fence="true" form="prefix">[</mml:mo>
							<mml:msub>
								<mml:mi>s</mml:mi>
								<mml:mrow>
									<mml:mi>q</mml:mi>
								</mml:mrow>
							</mml:msub>
							<mml:mo>≥</mml:mo>
							<mml:msub>
								<mml:mi>τ</mml:mi>
								<mml:mrow>
									<mml:mi>A</mml:mi>
								</mml:mrow>
							</mml:msub>
							<mml:mo stretchy="true" fence="true" form="postfix">]</mml:mo>
						</mml:mrow>
						<mml:mi>·</mml:mi>
						<mml:mrow>
							<mml:mn mathvariant="bold">1</mml:mn>
						</mml:mrow>
						<mml:mrow>
							<mml:mo stretchy="true" fence="true" form="prefix">[</mml:mo>
							<mml:mo>cov</mml:mo>
							<mml:mrow>
								<mml:mo stretchy="true" fence="true" form="prefix">(</mml:mo>
								<mml:msub>
									<mml:mi>b</mml:mi>
									<mml:mrow>
										<mml:mi>q</mml:mi>
									</mml:mrow>
								</mml:msub>
								<mml:mo>,</mml:mo>
								<mml:msub>
									<mml:mi>b</mml:mi>
									<mml:mrow>
										<mml:mi>w</mml:mi>
									</mml:mrow>
								</mml:msub>
								<mml:mo stretchy="true" fence="true" form="postfix">)</mml:mo>
							</mml:mrow>
							<mml:mo>≥</mml:mo>
							<mml:msub>
								<mml:mi>θ</mml:mi>
								<mml:mrow>
									<mml:mi>A</mml:mi>
								</mml:mrow>
							</mml:msub>
							<mml:mo stretchy="true" fence="true" form="postfix">]</mml:mo>
						</mml:mrow>
						<mml:mi>·</mml:mi>
						<mml:msub>
							<mml:mo>max</mml:mo>
							<mml:mrow>
								<mml:mi>k</mml:mi>
								<mml:mo>∈</mml:mo>
								<mml:msubsup>
									<mml:mi>P</mml:mi>
									<mml:mrow>
										<mml:mi>t</mml:mi>
									</mml:mrow>
									<mml:mrow>
										<mml:mi>T</mml:mi>
									</mml:mrow>
								</mml:msubsup>
							</mml:mrow>
						</mml:msub>
						<mml:mrow>
							<mml:mn mathvariant="bold">1</mml:mn>
						</mml:mrow>
						<mml:mrow>
							<mml:mo stretchy="true" fence="true" form="prefix">[</mml:mo>
							<mml:mo>cov</mml:mo>
							<mml:mrow>
								<mml:mo stretchy="true" fence="true" form="prefix">(</mml:mo>
								<mml:msub>
									<mml:mi>b</mml:mi>
									<mml:mrow>
										<mml:mi>q</mml:mi>
									</mml:mrow>
								</mml:msub>
								<mml:mo>,</mml:mo>
								<mml:msub>
									<mml:mi>b</mml:mi>
									<mml:mrow>
										<mml:mi>k</mml:mi>
									</mml:mrow>
								</mml:msub>
								<mml:mo stretchy="true" fence="true" form="postfix">)</mml:mo>
							</mml:mrow>
							<mml:mo>≥</mml:mo>
							<mml:msub>
								<mml:mi>θ</mml:mi>
								<mml:mrow>
									<mml:mtext>bed </mml:mtext>
								</mml:mrow>
							</mml:msub>
							<mml:mo stretchy="true" fence="true" form="postfix">]</mml:mo>
						</mml:mrow>
						<mml:mo stretchy="true" fence="true" form="postfix">)</mml:mo>
					</mml:mrow>
				</mml:mrow>
			</mml:math>
			<p>where τ and θ are again confidence and coverage thresholds; a value of one marks a procedural violation associated with the worker at that frame. These indicators are combined into an instantaneous safety-violation factor for the worker at frame t:</p>
			<mml:math display="inline">
				<mml:mrow>
					<mml:msub>
						<mml:mi>φ</mml:mi>
						<mml:mrow>
							<mml:mi>t</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo stretchy="false">(</mml:mo>
					<mml:mi>w</mml:mi>
					<mml:mo stretchy="false">)</mml:mo>
					<mml:mo>=</mml:mo>
					<mml:msub>
						<mml:mi>λ</mml:mi>
						<mml:mrow>
							<mml:mi>H</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mrow>
						<mml:mo stretchy="true" fence="true" form="prefix">(</mml:mo>
						<mml:mn>1</mml:mn>
						<mml:mo>−</mml:mo>
						<mml:msub>
							<mml:mi>h</mml:mi>
							<mml:mrow>
								<mml:mi>t</mml:mi>
							</mml:mrow>
						</mml:msub>
						<mml:mo stretchy="false">(</mml:mo>
						<mml:mi>w</mml:mi>
						<mml:mo stretchy="false">)</mml:mo>
						<mml:mo stretchy="true" fence="true" form="postfix">)</mml:mo>
					</mml:mrow>
					<mml:mo>+</mml:mo>
					<mml:msub>
						<mml:mi>λ</mml:mi>
						<mml:mrow>
							<mml:mi>V</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mrow>
						<mml:mo stretchy="true" fence="true" form="prefix">(</mml:mo>
						<mml:mn>1</mml:mn>
						<mml:mo>−</mml:mo>
						<mml:msub>
							<mml:mi>v</mml:mi>
							<mml:mrow>
								<mml:mi>t</mml:mi>
							</mml:mrow>
						</mml:msub>
						<mml:mo stretchy="false">(</mml:mo>
						<mml:mi>w</mml:mi>
						<mml:mo stretchy="false">)</mml:mo>
						<mml:mo stretchy="true" fence="true" form="postfix">)</mml:mo>
					</mml:mrow>
					<mml:mo>+</mml:mo>
					<mml:msub>
						<mml:mi>λ</mml:mi>
						<mml:mrow>
							<mml:mi>A</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:msub>
						<mml:mi>n</mml:mi>
						<mml:mrow>
							<mml:mi>t</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo stretchy="false">(</mml:mo>
					<mml:mi>w</mml:mi>
					<mml:mo stretchy="false">)</mml:mo>
					<mml:mo>,</mml:mo>
					<mml:mspace width="0.278em"/>
					<mml:msub>
						<mml:mi>λ</mml:mi>
						<mml:mrow>
							<mml:mi>H</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo>+</mml:mo>
					<mml:msub>
						<mml:mi>λ</mml:mi>
						<mml:mrow>
							<mml:mi>V</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo>+</mml:mo>
					<mml:msub>
						<mml:mi>λ</mml:mi>
						<mml:mrow>
							<mml:mi>A</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo>=</mml:mo>
					<mml:mn>1</mml:mn>
				</mml:mrow>
			</mml:math>
			<p>in which the non-negative severity weights λ express the relative danger of a missing helmet, a missing vest and an unsafe climb, normalized to sum to one, so that the factor lies between zero and one. It is therefore an interpretable, graded measure: zero when the worker is fully compliant, and approaching one as several severe violations occur together.</p>
			<p>A decision taken on a single frame would inherit the instability of the detector, whose safety-critical classes attain only moderate confidence, as the per-class metrics show (Table 2). The track produced by the Kalman filter links the detections of one worker across successive frames and so lets the evidence be accumulated rather than judged frame by frame. The factor is aggregated along the track by an exponential moving average, written in the same recursive form as the filter itself:</p>
			<mml:math display="inline">
				<mml:mrow>
					<mml:msub>
						<mml:mi>Φ</mml:mi>
						<mml:mrow>
							<mml:mi>t</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo stretchy="false">(</mml:mo>
					<mml:mi>w</mml:mi>
					<mml:mo stretchy="false">)</mml:mo>
					<mml:mo>=</mml:mo>
					<mml:mi>γ</mml:mi>
					<mml:msub>
						<mml:mi>Φ</mml:mi>
						<mml:mrow>
							<mml:mi>t</mml:mi>
							<mml:mo>−</mml:mo>
							<mml:mn>1</mml:mn>
						</mml:mrow>
					</mml:msub>
					<mml:mo stretchy="false">(</mml:mo>
					<mml:mi>w</mml:mi>
					<mml:mo stretchy="false">)</mml:mo>
					<mml:mo>+</mml:mo>
					<mml:mo stretchy="false">(</mml:mo>
					<mml:mn>1</mml:mn>
					<mml:mo>−</mml:mo>
					<mml:mi>γ</mml:mi>
					<mml:mo stretchy="false">)</mml:mo>
					<mml:msub>
						<mml:mi>φ</mml:mi>
						<mml:mrow>
							<mml:mi>t</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo stretchy="false">(</mml:mo>
					<mml:mi>w</mml:mi>
					<mml:mo stretchy="false">)</mml:mo>
					<mml:mo>,</mml:mo>
					<mml:mspace width="0.278em"/>
					<mml:mi>γ</mml:mi>
					<mml:mo>∈</mml:mo>
					<mml:mo stretchy="false">[</mml:mo>
					<mml:mn>0</mml:mn>
					<mml:mo>,</mml:mo>
					<mml:mn>1</mml:mn>
					<mml:mo stretchy="false">)</mml:mo>
				</mml:mrow>
			</mml:math>
			<p>where the smoothing coefficient γ governs the memory of the estimate: larger values suppress brief false detections more strongly but react more slowly to a genuine violation. A violation is registered when the aggregated factor exceeds a decision threshold:</p>
			<mml:math display="inline">
				<mml:mrow>
					<mml:msub>
						<mml:mi>V</mml:mi>
						<mml:mrow>
							<mml:mi>t</mml:mi>
						</mml:mrow>
					</mml:msub>
					<mml:mo stretchy="false">(</mml:mo>
					<mml:mi>w</mml:mi>
					<mml:mo stretchy="false">)</mml:mo>
					<mml:mo>=</mml:mo>
					<mml:mrow>
						<mml:mn mathvariant="bold">1</mml:mn>
					</mml:mrow>
					<mml:mrow>
						<mml:mo stretchy="true" fence="true" form="prefix">[</mml:mo>
						<mml:msub>
							<mml:mi>Φ</mml:mi>
							<mml:mrow>
								<mml:mi>t</mml:mi>
							</mml:mrow>
						</mml:msub>
						<mml:mo stretchy="false">(</mml:mo>
						<mml:mi>w</mml:mi>
						<mml:mo stretchy="false">)</mml:mo>
						<mml:mo>≥</mml:mo>
						<mml:mi>Θ</mml:mi>
						<mml:mo stretchy="true" fence="true" form="postfix">]</mml:mo>
					</mml:mrow>
				</mml:mrow>
			</mml:math>
			<p>optionally with the additional condition that this hold over a minimum number of consecutive frames, which removes residual flicker.</p>
			<p>The formulation makes the behavior of the system explicit and reproducible. Its only free parameters — the confidence and coverage thresholds, the severity weights, the smoothing coefficient γ and the decision threshold Θ — each carry a clear operational meaning, and in this work they were fixed on the validation subset: the confidence thresholds at the values that maximized the per-class F-measure, and the smoothing and decision thresholds so as to balance missed violations against false alarms on the held-out recordings. Because the factor is a single scalar computed from quantities that the detector already produces, the same decision rule transfers without modification to a stronger detector; only the confidence thresholds need to be recalibrated.</p>
			<p>4. Experimental evaluation</p>
			<p>The detector’s quality was assessed using metrics common in detection tasks. Precision indicates the proportion of correct detections among all detections, recall indicates the proportion of detected violations among all those present, and their generalization is the mean average precision (mAP), computed at a box-overlap threshold of IoU = 0.5 (mAP@0.5) and averaged over the range of thresholds 0.5…0.95 (mAP@0.5:0.95). The system’s suitability for real-time operation is characterized by the number of frames processed per second (FPS). The values of these metrics by class are given in Table 2.</p>
			<table-wrap id="T2">
				<label>Table 2</label>
				<caption>
					<p>Detection quality metrics by class</p>
				</caption>
				<table>
					<tr>
						<td>Class</td>
						<td>Precision</td>
						<td>Recall</td>
						<td>mAP@0.5</td>
					</tr>
					<tr>
						<td>Helmet (helmet)</td>
						<td>0.79</td>
						<td>0.68</td>
						<td>0.72</td>
					</tr>
					<tr>
						<td>Vest (vest)</td>
						<td>0.77</td>
						<td>0.66</td>
						<td>0.70</td>
					</tr>
					<tr>
						<td>Truck (kamaz)</td>
						<td>0.92</td>
						<td>0.89</td>
						<td>0.91</td>
					</tr>
					<tr>
						<td>Climbing without a ladder (no_stairs)</td>
						<td>0.71</td>
						<td>0.58</td>
						<td>0.63</td>
					</tr>
					<tr>
						<td>All classes</td>
						<td>0.80</td>
						<td>0.70</td>
						<td>0.74</td>
					</tr>
				</table>
			</table-wrap>
			<p>The system’s performance was tested on video recordings made in an operating production workshop. An example of a recognized scene is shown in Figure 1: the system highlights workers wearing high-visibility vests and helmets, identifies the truck, and interprets climbing into its bed without a step ladder as a violation.</p>
			<fig id="F1">
				<label>Figure 1</label>
				<caption>
					<p>Detection of a violation: climbing into a KamAZ truck bed without a ladder</p>
				</caption>
				<alt-text>Detection of a violation: climbing into a KamAZ truck bed without a ladder</alt-text>
				<graphic ns1:href="/media/images/2026-08-12/6df766f5-4146-4ebd-a830-ee2e31274822.png"/>
			</fig>
			<p>The low confidence values for the classes that are central to occupational safety point to the main limitation of the current version — the moderate size of the training set (on the order of a thousand images) and the comparatively low input resolution (416 pixels). Both restrict the network’s ability to distinguish small details of equipment in a complex, object-dense scene. Improvements in quality can be expected from expanding and balancing the dataset across classes, increasing the input resolution, and applying more capacious network configurations while maintaining acceptable performance.</p>
			<p>Object tracking with a Kalman filter smooths the results across frames and reduces the influence of individual missed detections: even if a helmet or vest is not recognized in a particular frame, a stable worker track prevents the false registration of a violation, while the accumulation of evidence over a series of frames, conversely, increases the reliability of recording an actual violation.</p>
			<p>Having recorded a violation, the system automatically generates a report that includes the frame from the moment of the violation and its type — the absence of a helmet or vest, climbing into a truck bed without a ladder, and the like. The result of recognition thus becomes not merely an annotated frame but a document suitable for further review, which can be used by the enterprise’s occupational-safety service.</p>
			<p>When working with a video stream, processing is expected in a near-real-time mode, with an average performance on the order of 30 frames per second.</p>
			<p>The moderate accuracy of the safety-critical classes reported above is in part a property of the YOLOv5 detector adopted for this study, and it is therefore useful to set that choice against the current state of the art. Table 3 collects the detection accuracy reported in the literature for the original YOLOv5 and for two recent detectors — the attention-centric YOLO12 and the end-to-end YOLO26 — on the public COCO benchmark.</p>
			<table-wrap id="T3">
				<label>Table 3</label>
				<caption>
					<p>Detection accuracy reported in the literature on the COCO val2017 benchmark</p>
				</caption>
				<table>
					<tr>
						<td>Detector</td>
						<td>Year</td>
						<td>Detection paradigm</td>
						<td>mAP@0.5:0.95, %</td>
					</tr>
					<tr>
						<td>YOLOv5-s</td>
						<td>2020</td>
						<td>anchor-based CNN, NMS</td>
						<td>37.4</td>
					</tr>
					<tr>
						<td>YOLOv5-m</td>
						<td>2020</td>
						<td>anchor-based CNN, NMS</td>
						<td>45.4</td>
					</tr>
					<tr>
						<td>YOLO12-s</td>
						<td>2025</td>
						<td>attention-centric, NMS</td>
						<td>48.0</td>
					</tr>
					<tr>
						<td>YOLO12-m</td>
						<td>2025</td>
						<td>attention-centric, NMS</td>
						<td>52.5</td>
					</tr>
					<tr>
						<td>YOLO26-s</td>
						<td>2025</td>
						<td>end-to-end, NMS-free</td>
						<td>≈47.2</td>
					</tr>
					<tr>
						<td>YOLO26-m</td>
						<td>2025</td>
						<td>end-to-end, NMS-free</td>
						<td>≈51.5</td>
					</tr>
				</table>
			</table-wrap>
			<p>Two observations follow. First, the figures in Table 3 are not directly comparable with those of Table 2: they are obtained on a different dataset of eighty general-purpose classes and at a stricter, averaged-IoU threshold, whereas Table 2 reports mAP@0.5 on roughly a thousand domain-specific images. The comparison, therefore, concerns the relative capability of the architectures, not the accuracy attainable on the present task. Second, within that comparison, the advantage of the newer models is large and systematic: at a comparable model size the small YOLO12 model exceeds the original YOLOv5 baseline by more than 10 percentage points of mAP (48.0 against 37.4), and the small and medium YOLO26 models reach a similar level while additionally dispensing with non-maximum suppression and with the distribution focal loss and reducing processor inference time by up to 43% </p>
			<p>[7][8]</p>
			<p>These properties bear directly on the limitation identified earlier. The weak classes in this study are the small protective items, and the small-target-aware training of YOLO26 together with the higher accuracy of YOLO12 address exactly that weakness, while the faster, suppression-free inference of YOLO26 suits the near-real-time, edge-oriented deployment intended here. Replacing the detector with one of these architectures and re-evaluating it on the enterprise dataset under the protocol of Table 2 is consequently the most promising next step; it is left to future work because it requires re-annotation and re-training on the target classes rather than the reuse of published weights. It must be stressed that the values in Table 3 are reproduced from the literature and were not measured on the data of this study, and that the decision model formalized above is independent of the detector and would carry over unchanged to any of these networks.</p>
			<p>5. Conclusion</p>
			<p>This work investigated computer vision methods as applied to the detection of occupational-safety violations and implemented an information system that analyzes the video stream from cameras, monitors safety compliance, and compiles violations into reports. Its core is the YOLOv5 detector model, fine-tuned for the needs of a specific enterprise </p>
			<p>[23]</p>
			<p>The principal contribution of the work is a formalized model of the safety-violation factor — an interpretable scalar that maps the detector outputs into a graded, track-aggregated decision through explicit spatial and temporal criteria — which replaces ad hoc heuristics with a reproducible rule governed by a small set of tunable parameters. It is this model that allows the system to recognize not only the presence of personal protective equipment but also a dangerous action defined by the spatial context of the scene. A further distinctive feature is the integrity of the technological pipeline: versioned data preparation, detector fine-tuning, track-level violation scoring with a Kalman filter, and the automatic logging of violations.</p>
			<p>The results obtained confirm the fundamental viability of the approach and, at the same time, outline directions for its development. The main reserve for improving quality is associated with increasing the volume and diversity of the annotated data, raising the input resolution, and fine-tuning the decision thresholds; the addition of new violation classes and the integration of the system with the alerting facilities already in operation at the enterprise also appear promising. A comparison with current architectures indicates that the clearest route to higher accuracy on the small, safety-critical classes is to replace the YOLOv5 detector with a more recent network — the attention-centric YOLO12 or the end-to-end, small-target-oriented YOLO26 — and to re-evaluate it on the enterprise dataset; since the proposed safety-violation factor is defined over generic detector outputs, such a substitution leaves the decision model itself unchanged.</p>
			<p>The practical significance of the work is that the proposed solution can be embedded into existing video-surveillance infrastructure and applied across a wide variety of industries — from metalworking and construction to transport and warehouse logistics — where compliance with occupational-safety requirements is critically important.</p>
		</sec>
		<sec sec-type="supplementary-material">
			<title>Additional File</title>
			<p>The additional file for this article can be found as follows:</p>
			<supplementary-material xmlns:xlink="http://www.w3.org/1999/xlink" id="S1" xlink:href="https://doi.org/10.5334/cpsy.78.s1">
				<!--[<inline-supplementary-material xlink:title="local_file" xlink:href="https://research-journal.org/media/articles/26006.docx">26006.docx</inline-supplementary-material>]-->
				<!--[<inline-supplementary-material xlink:title="local_file" xlink:href="https://research-journal.org/media/articles/26006.pdf">26006.pdf</inline-supplementary-material>]-->
				<label>Online Supplementary Material</label>
				<caption>
					<p>
						Further description of analytic pipeline and patient demographic information. DOI:
						<italic>
							<uri>https://doi.org/10.60797/IRJ.2026.170.51</uri>
						</italic>
					</p>
				</caption>
			</supplementary-material>
		</sec>
	</body>
	<back>
		<ack>
			<title>Acknowledgements</title>
			<p/>
		</ack>
		<sec>
			<title>Competing Interests</title>
			<p/>
		</sec>
		<ref-list>
			<ref id="B1">
				<label>1</label>
				<mixed-citation publication-type="confproc">Girshick R. Fast R-CNN / R. Girshick // 2015 IEEE International Conference on Computer Vision (ICCV). — 2015. — P. 1440–1448. — DOI: 10.1109/ICCV.2015.169.</mixed-citation>
			</ref>
			<ref id="B2">
				<label>2</label>
				<mixed-citation publication-type="confproc">Ren S. Faster R-CNN: Towards real-time object detection with region proposal networks / S. Ren, K. He, R. Girshick [et al.] // IEEE Transactions on Pattern Analysis and Machine Intelligence. — 2017. — № 39 (6). — P. 1137–1149. — DOI: 10.1109/TPAMI.2016.2577031.</mixed-citation>
			</ref>
			<ref id="B3">
				<label>3</label>
				<mixed-citation publication-type="confproc">Redmon J. You only look once: Unified, real-time object detection / J. Redmon, S. Divvala, R. Girshick [et al.] // 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) — 2016. — P. 779–788. — DOI: 10.1109/CVPR.2016.91.</mixed-citation>
			</ref>
			<ref id="B4">
				<label>4</label>
				<mixed-citation publication-type="confproc">Redmon J. YOLO9000: Better, faster, stronger / J. Redmon, A. Farhadi // 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). — 2017. — P. 6517–6525. — DOI: 10.1109/CVPR.2017.690.</mixed-citation>
			</ref>
			<ref id="B5">
				<label>5</label>
				<mixed-citation publication-type="confproc">Redmon J. YOLOv3: An incremental improvement / J. Redmon, A. Farhadi // arXiv. — 2018. — DOI: 10.48550/arXiv.1804.02767.</mixed-citation>
			</ref>
			<ref id="B6">
				<label>6</label>
				<mixed-citation publication-type="confproc">Wang Z. Fast personal protective equipment detection for real construction sites using deep learning approaches / Z. Wang, Y. Wu, L. Yang [et al.] // Sensors. — 2021. — № 21 (10). — P. 3478. — DOI: 10.3390/s21103478.</mixed-citation>
			</ref>
			<ref id="B7">
				<label>7</label>
				<mixed-citation publication-type="confproc">Tian Y. YOLOv12: Attention-centric real-time object detectors / Y. Tian, Q. Ye, D. Doermann // arXiv. — 2025. — DOI: 10.48550/arXiv.2502.12524.</mixed-citation>
			</ref>
			<ref id="B8">
				<label>8</label>
				<mixed-citation publication-type="confproc">Kalfaoglu M.E. Ultralytics YOLO26: Unified real-time end-to-end vision models / M.E. Kalfaoglu // arXiv. — 2026. — DOI: 10.48550/arXiv.2606.03748.</mixed-citation>
			</ref>
			<ref id="B9">
				<label>9</label>
				<mixed-citation publication-type="confproc">Lin T.-Y. Feature pyramid networks for object detection / T.-Y. Lin, P. Dollár, R. Girshick [et al.] // 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). — 2017. — P. 936–944. IEEE. — DOI: 10.1109/CVPR.2017.106.</mixed-citation>
			</ref>
			<ref id="B10">
				<label>10</label>
				<mixed-citation publication-type="confproc">Bodla N. Soft-NMS — Improving object detection with one line of code / N. Bodla, B. Singh, R. Chellappa [et al.] // 2017 IEEE International Conference on Computer Vision (ICCV) — 2017. — P. 5562–5570. — DOI: 10.1109/ICCV.2017.593.</mixed-citation>
			</ref>
			<ref id="B11">
				<label>11</label>
				<mixed-citation publication-type="confproc">Liu W. SSD: Single shot multibox detector / W. Liu, D. Anguelov, D. Erhan [et al.] // Computer Vision – ECCV 2016. — 2016. — P. 21–37. — DOI: 10.1007/978-3-319-46448-0_2.</mixed-citation>
			</ref>
			<ref id="B12">
				<label>12</label>
				<mixed-citation publication-type="confproc">Wang C.-Y. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors / C.-Y. Wang, A. Bochkovskiy, H.-Y.M. Liao // 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). — 2023. — P. 7464–7475. — DOI: 10.1109/CVPR52729.2023.00721.</mixed-citation>
			</ref>
			<ref id="B13">
				<label>13</label>
				<mixed-citation publication-type="confproc">Jocher G. YOLOv5 by Ultralytics (Version 7.0) / G. Jocher // Zenodo. — 2020. — DOI: 10.5281/zenodo.3908559.</mixed-citation>
			</ref>
			<ref id="B14">
				<label>14</label>
				<mixed-citation publication-type="confproc">Elfwing S. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning / S. Elfwing, E. Uchibe, K. Doya // Neural Networks. — 2018. — № 107. — P. 3–11. — DOI: 10.1016/j.neunet.2017.12.012.</mixed-citation>
			</ref>
			<ref id="B15">
				<label>15</label>
				<mixed-citation publication-type="confproc">Bochkovskiy A. YOLOv4: Optimal speed and accuracy of object detection / A. Bochkovskiy, C.-Y. Wang, H.-Y.M. Liao // arXiv. — 2020. — DOI: 10.48550/arXiv.2004.10934.</mixed-citation>
			</ref>
			<ref id="B16">
				<label>16</label>
				<mixed-citation publication-type="confproc">Lin T.-Y. Focal loss for dense object detection / T.-Y. Lin, P. Goyal, R. Girshick [et al.] // 2017 IEEE International Conference on Computer Vision (ICCV) — 2017. — P. 2980–2988. — DOI: 10.1109/ICCV.2017.324.</mixed-citation>
			</ref>
			<ref id="B17">
				<label>17</label>
				<mixed-citation publication-type="confproc">Rezatofighi H. Generalized intersection over union: A metric and a loss for bounding box regression / H. Rezatofighi, N. Tsoi, J. Gwak [et al.] // 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) — 2019. — P. 658–666. — DOI: 10.1109/CVPR.2019.00075.</mixed-citation>
			</ref>
			<ref id="B18">
				<label>18</label>
				<mixed-citation publication-type="confproc">Kalman R.E. A new approach to linear filtering and prediction problems / R.E. Kalman // Journal of Basic Engineering, — 1960. — № 82 (1). — P. 35–45. — DOI: 10.1115/1.3662552.</mixed-citation>
			</ref>
			<ref id="B19">
				<label>19</label>
				<mixed-citation publication-type="confproc">Terven J. A comprehensive review of YOLO architectures in computer vision: From YOLOv1 to YOLOv8 and YOLO-NAS / J. Terven, D.-M. Córdova-Esparza, J.-A. Romero-González // Machine Learning and Knowledge Extraction. — 2023. — № 5 (4). — P. 1680–1716. — DOI: 10.3390/make5040083.</mixed-citation>
			</ref>
			<ref id="B20">
				<label>20</label>
				<mixed-citation publication-type="confproc">Lin T.-Y. Microsoft COCO: Common objects in context / T.-Y. Lin, M. Maire, S. Belongie [et al.] // Computer Vision – ECCV 2014. — 2014. — P. 740–755. — DOI: 10.1007/978-3-319-10602-1_48.</mixed-citation>
			</ref>
			<ref id="B21">
				<label>21</label>
				<mixed-citation publication-type="confproc">Wang C.-Y. Scaled-YOLOv4: Scaling cross stage partial network / C.-Y. Wang, A. Bochkovskiy, H.-Y.M. Liao // 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). — 2021. — P. 13029–13038. — DOI: 10.1109/CVPR46437.2021.01283.</mixed-citation>
			</ref>
			<ref id="B22">
				<label>22</label>
				<mixed-citation publication-type="confproc">Hussain M. YOLO-v1 to YOLO-v8, the rise of YOLO and its complementary nature toward digital manufacturing and industrial defect detection / M. Hussain // Machines. — 2023. — № 11 (7). — P. 677. — DOI: 10.3390/machines11070677.</mixed-citation>
			</ref>
			<ref id="B23">
				<label>23</label>
				<mixed-citation publication-type="confproc">Nath N.D. Deep learning for site safety: Real-time detection of personal protective equipment / N.D. Nath, A.H. Behzadan, S.G. Paal // Automation in Construction, — 2020. — № 112. — Art. 103085. — DOI: 10.1016/j.autcon.2020.103085.</mixed-citation>
			</ref>
		</ref-list>
	</back>
	<fundings>
		<funding lang="RUS">Данная работа/публикация выполнена за счет гранта, предоставленного Академией наук Республики Татарстан образовательным организациям высшего образования, научным и иным организациям на поддержку планов развития кадрового потенциала в части стимулирования их научных и научно-педагогических работников к защите докторских диссертаций и выполнению научно-исследовательских работ (соглашение №15/2025-ПД-КАИ от 22.12.2025).</funding>
		<funding lang="ENG">This work/publication was funded by a grant from the Academy of Sciences of the Republic of Tatarstan provided to higher education institutions, scientific and other organizations to support human resource development plans in terms of encouraging their research and academic staff to defend doctoral dissertations and conduct research activities. (Agreement No. 15/2025-PD-KAI dated 22 December 2025).</funding>
	</fundings>
</article>