EC833OE: FUNDAMENTALSOFSOCIALNETWORKS
(OE-III)
[Link] L T P C
3 0 03
Course Objectives:
1. Togiveoverviewonsocialnetworks.
2. To make social media, information networks and World Wide Web concepts more
familiar.
3. To provide knowledge on social network ties.
4. To provide knowledge on power laws related to information networks.
Courseoutcomes:uponcompletingthiscoursethestudentswillbeableto
1. Understandconceptslikesmall-
worldexperimentandsnowballsamplingrelatedtosocial networks.
2. Get knowledge on ties, weak ties and their strength.
3. Know about structure of the web, modern web search,linkanalysisusinghubs.
4. AcquireknowledgeonpowerlawsandanalysisofRich-get-Richerphenomena.
Course PO1 PO2 PO3 PO4 PO5 PO6 PO7 PO8 PO9 PO10 PO11 PO12
CO1 1 1 - 1 - - - - 1 - - 1
CO2 1 1 - 1 - - - - 1 - - 1
CO3 1 1 1 1 - - - - 1 - - 1
CO4 1 1 1 1 - - - - 1 - - 1
UNIT - I: Introduction to social networks: The Empirical Study of Social
Networks, Interviews and Questionnaires, Direct Observation, Data from Archival
or Third-Party Records, Affiliation Networks, The Small-World Experiment,
Snowball Sampling, Contact Tracing, and Random Walks.
UNIT-
II:GraphtheoryandSocialNetworks:Basicdefinitions,PathsandConnectivity,Thestr
ength of weak ties, Tie Strength and Network Structure in Large-Scale Data, Tie
strength, social media, passive engagement.
UNIT-III:
InformationnetworksandWorldWideWeb:TheWorldWideWeb,InformationNetworks,
Hypertext,andAssociativeMemory,TheWebasaDirectedGraph,TheBow-TieStructureoftheWeb,
theemergenceofweb2.0,SearchingtheWeb:TheProblemofRankingLinkAnalysisusingHubsand
Authorities, PageRank, Applying Link Analysis in Modern Web Search.
UNIT - IV: Power Laws and Rich-Get-Richer Phenomena: Popularity as a Network
Phenomenon,PowerLaws,Rich-Get-RicherModels,TheUnpredictabilityofRich-
GetRicherEffects,TheLongTail,TheEffectofSearchToolsandRecommendationSystems,AdvancedMat
erial:AnalysisofRich-Get- Richer Processes.
UNIT - V: The Small-World Phenomenon: Six Degrees of Separation, Structure and Randomness,
Decentralized Search, Modeling the Process of Decentralized Search, Empirical Analysis and
Generalized Models, Core-Periphery Structures and Difficulties in Decentralized Search, Advanced
Material: Analysis of Decentralized Search.
TEXTBOOKS:
1. [Link]“Networksanintroduction”OxfordUniversityPress2010.
2. Networks,CrowdsandMarketsbyDavidEasleyandJonKleinberg,CambridgeUn
iversity Press, 2010.
REFERENCEBOOKS:
1. [Link],PrincetonUniversityPress,2010.
2. MaksimTsvetovatandAlexanderKouznetsov.“SocialNetworkAnalysisforStart
ups”.O’Reilly Media, 2011.
UNIT – I
INTRODUCTION TO SOCIAL NETWORKS
1. Social Networks – An Introduction
A social network is a set of social actors (individuals, organizations, communities, or
systems) connected through relationships, called edges. These relationships may represent
friendship, professional contact, information sharing, communication, kinship, trade, or
influence. The study of social networks focuses on how these relationships shape behavior,
opportunities, knowledge flow, and social structure.
Social networks can be represented mathematically using graphs, where nodes represent actors
and edges represent connections. Social Network Analysis (SNA) applies graph theory, statistics,
and computational techniques to understand network properties.
Fundamentals of social networks involve nodes (people/entities) and edges (relationships)
forming structures, analyzed through Social Network Analysis (SNA) using graph theory to
understand patterns like clusters (communities) and influential figures (centrality), all while
managing core functions like profile creation, content sharing, and instant communication, but
facing challenges in data quality, privacy, and scalability, as they serve purposes from personal
connection to business marketing
Core Concepts
Nodes (Actors/Entities): Individuals, groups, or organizations in the network.
Edges (Ties/Links): Relationships connecting nodes, such as friendship, family, professional, or
even antagonistic ties.
Sociogram: A visual graph representing the network structure with dots (nodes) and lines
(edges).
Sociomatrix: A data structure (matrix) storing network relationships for analysis.
Key Features & Functions
Profile Creation: Users build personal digital identities.
Content Sharing: Sharing updates, photos, videos.
Connection: Adding friends, following, joining groups.
Communication: Real-time messaging, chat, video calls.
Information Flow: Networks facilitate rapid spread of news and trends.
Social Network Analysis (SNA) Principles
Graph Theory: Using mathematical structures to study network properties.
Centrality: Identifying key nodes (influencers, hubs).
Homophily: The tendency for individuals to connect with similar others.
Communities/Clusters: Groups of densely connected nodes.
Network Dynamics: How networks change over time (friendships forming/dissolving).
Purposes & Applications
Personal: Staying connected with family/friends, belonging, social recognition.
Business: Marketing, brand building, customer retention.
Research: Studying social behavior, influence, resource mobilization.
Challenges
Privacy & Security: Protecting user data.
Scalability: Handling massive amounts of data.
Data Quality: Ensuring consistency across diverse sources.
User Awareness: Making complex tech accessible.
2. Empirical Study of Social Networks
The empirical study of social networks refers to the systematic collection and analysis of real-
world data to understand social interactions. Instead of theoretical assumptions, empirical studies
rely on observed or recorded relationships.
Objectives:
Discover interaction patterns
Identify influential individuals
Study diffusion of information and diseases
Understand formation of communities and social groups
Support decision making in marketing, health, education, and governance
Empirical network studies have become more prominent due to the availability of digital trace
data from social media platforms
3. Methods of Data Collection
3.1 Interviews and Questionnaires
In this method, individuals are directly asked to report their social connections, frequency of
communication, and type of relationships.
Advantages:
Accurate and detailed information
Allows collection of qualitative insights
Suitable for small, closed populations
Disadvantages:
Time-consuming
Subject to recall bias
Not suitable for large-scale networks
3.2 Direct Observation
Here, the researcher directly observes interactions among individuals in real-life settings such as
classrooms, workplaces, hospitals, and markets.
Used to study:
Participation behavior
Group communication
Leadership patterns
It provides highly realistic data but requires extensive human effort.
3.3 Data from Archival or Third-Party Records
This method uses already available records such as:
Phone call logs
=-0985dfc vb/\7/Email archives
Social media logs
Organizational and government databases
It enables large-scale network analysis involving millions of nodes and edges.
4. Affiliation Networks
Affiliation networks in social networks map relationships between two distinct sets of entities,
like people and the groups/events they join, using bipartite graphs where links only cross
between the sets (e.g., person to club, not person to person directly). They reveal hidden social
ties by showing how shared memberships (e.g., common clubs, co-authored papers) connect
individuals, allowing for analysis of social structures and relationships through techniques like
'folding' the graph to create a one-mode network of just people.
Characteristics
Two Types of Nodes: One set for actors (people, authors) and another for affiliations (groups,
events, organizations).
Bipartite Structure: Ties (edges) exist only between an actor and an affiliation, never between
two actors or two affiliations directly.
Hidden Ties: Shared affiliations create implicit connections between actors, forming a "folded"
social network where co-membership implies potential friendship or acquaintance.
Applications
Understanding influence and structure within communities.
Analyzing collaboration patterns (e.g., co-authorship networks).
Mapping relationships in fields like counter-terrorism, linking individuals through shared events
or organizations.
5. The Small-World Experiment
Small world experiments in social network analysis (SNA) demonstrate that any two people in a
large network are connected by surprisingly short chains of acquaintances, famously illustrated
by the "six degrees of separation" concept, with early studies by Milgram showing an average of
6 steps, later confirmed by Watts-Strogatz modeling and modern digital studies
(like Facebook's), revealing robust, interconnected structures with high clustering and short
paths, crucial for understanding information spread.
Key Experiments & Concepts:
1. Milgram's "Small World" Experiment (1960s):
1. Method: Sent letters from people in Nebraska to a target in Boston, asking senders to forward the
letter to someone they knew on a first-name basis, hoping to reach the target
2. Findings: Completed chains averaged around 5.9 steps (about six), popularizing the "six
degrees of separation" idea.
3. Significance: Showed that people could navigate vast networks using only local knowledge,
finding short paths efficiently.
2. Watts-Strogatz Model (1998):
0. Contribution: Provided a mathematical framework, showing how networks can have high
clustering (like regular grids) and short average path lengths (like random networks) by adding a
few random "shortcuts" (rewiring).
1. Key Properties: Characterizes small-world networks by both high clustering coefficient (tight-
knit groups) and low average path length (quick connections).
3. Modern Digital Experiments (e.g., Face book):
0. Method: Analyzed vast digital networks, like Facebook's, to find path lengths between users.
1. Findings: Confirmed the small-world phenomenon, with Facebook finding even shorter paths
(around 4.74 steps in 2011) than Milgram's original studies, showing increasing
interconnectedness.
Importance:
Explains rapid spread of information
Demonstrates social closeness
Used in marketing, epidemiology, and social media analysis
6. Snowball Sampling
Snowball sampling is a non-probability sampling technique where existing participants recruit
future participants. It is useful for studying hidden or hard-to-reach populations.
Snowball sampling (or chain-referral sampling) is a non-probability sampling technique where
researchers find initial participants, who then refer other people they know from their social
networks who fit the study's criteria, creating a growing sample like a rolling snowball. It's ideal
for studying hidden, rare, or hard-to-reach populations (like drug users, rare disease patients, or
members of niche groups) because it leverages existing trust and social connections to access
otherwise inaccessible communities. While efficient and cost-effective for these groups, its main
drawback is that it's not random, leading to potential biases as the sample reflects the
interconnectedness of the network rather than the general population.
How it Works:
1. Identify Seeds: Researchers find one or a few initial individuals (seeds) who meet the study's
criteria.
2. Ask for Referrals: These participants are asked to recommend other people they know who also
fit the criteria.
3. Chain Referral: Those referred individuals are then asked to refer more people, continuing the
chain.
4. Saturation: The process stops when the desired sample size is reached or when no new
participants can be found (saturation).
When to Use It:
Studying rare diseases or genetic disorders.
Researching marginalized or stigmatized groups (e.g., homeless individuals, illegal immigrants).
Understanding specific niche communities or subcultures.
Advantages
Access: Excellent for hard-to-reach populations.
Cost-Effective & Efficient: Can rapidly build samples in diverse locations.
Builds Trust: Leverages existing social trust for sensitive topics.
Disadvantages
Non-Random: Not statistically generalizable to the whole population.
Bias: Sample may overrepresent highly connected individuals or specific network clusters.
7. Contact Tracing
Contact tracing is the process of identifying and monitoring people who have come in contact
with an infected person to prevent disease spread.
or
Contact tracing is a public health method to find people exposed to an infectious disease
(contacts) from a known infected person (case) to stop further spread by isolating the infected
and quarantining the exposed, breaking transmission chains and preventing new outbreaks, vital
for diseases like COVID-19.
How it works:
1. Identify the Case:
A person tests positive for a contagious disease.
2. Interview the Case:
A health worker asks the infected person about places they visited and people they met (contacts)
while contagious.
3. Locate Contacts:
The health worker identifies potential contacts (close friends, family, coworkers, public places
visited).
4. Notify & Educate Contacts:
Contacts are informed they were exposed and advised to self-monitor for symptoms, test if
needed, and quarantine.
5. Monitor & Support:
Public health teams track symptoms and support contacts to ensure isolation, preventing onward
spread.
Purpose:
Break Chains of Transmission: Stop the virus/bacteria from moving from person to person.
Prevent New Infections: Reduce overall community spread.
Inform Public Health: Understand where and how the disease is spreading.
Tools:
Manual Interviews: Traditional method by health workers.
Digital Tools: Apps and location data (like during COVID-19) to help find contacts quickly.
Contact tracing is a core strategy in controlling infectious disease outbreaks, relying on
cooperation and confidentiality to protect communities.
Widely used during the COVID-19 pandemic.
8. Random Walks
A random walk is a stochastic process that randomly moves from one node to another.
A random walk on a graph is a stochastic process where a token moves between vertices (nodes)
by randomly selecting a neighbor at each step, forming a Markov chain. This process is
fundamental for analyzing graph connectivity, mixing times, and finding stationary distributions,
often described by the transition matrix (normalized adjacency matrix) and related to graph
properties like eigenvalues, with applications in search engines (Google's PageRank), network
analysis, and modeling physical processes.
How It Works
Starting Point: Begin at a specific vertex (node).
Movement: At each step, move to an adjacent vertex chosen uniformly at random (or
proportional to edge weights in a weighted graph).
Markovian Property: The next step only depends on the current vertex, not the entire path
taken so far.
Applications:
PageRank algorithm
Network exploration
Recommendation systems
Community detection